Architecture decision
Cloud vs. On-Premise vs. Hybrid DDoS Protection: Cost, Latency and Sovereignty
Last updated: August 2026 · Three architectures compared · Reading time ~16 min

Choose cloud scrubbing when your attack exposure is volumetric, your traffic is not regulated, and you cannot invest in hardware. Choose an on-premise appliance when latency matters, when application-layer attacks are your real threat, or when regulation makes cross-border inspection a compliance problem. Choose hybrid — an always-on inline appliance plus an upstream tier that engages only above your uplink capacity — when you need both, which in 2026 is most organisations above a few hundred Mbps of real traffic. The decision is not about which tier is stronger; it is about where the ceiling sits and whose law applies to the inspection. The inline half of a hybrid is filled by appliances carrying full L3–L7 coverage in a single device.
Almost every DDoS conversation starts in the wrong place. It starts with “which product is strongest”, when the decision that actually determines the outcome is architectural: where the inspection happens, who owns the equipment doing it, and which legal regime governs the traffic while it is being inspected. Get that wrong and the strongest product on the market will not save you — it will simply be strong in the wrong location.
This guide compares the three architectures on the four dimensions that decide real purchases: how fast each one actually mitigates, where each one hits a hard ceiling, what each one costs across five years, and whose jurisdiction inspects your users’ traffic.

The three architectures
The vocabulary is muddier than it should be, so it is worth fixing the terms before comparing them.
Cloud scrubbing routes your traffic through a provider’s global network. Attack traffic is filtered in their scrubbing centres and clean traffic is forwarded to you, usually over a tunnel. Engagement is either always-on — all traffic flows through the provider permanently — or on-demand, where you divert only when you detect an attack, typically by announcing your prefixes through the provider with BGP.
On-premise mitigation places a purpose-built appliance at your network edge, inline between your border router and your services. It inspects every packet at line rate, holds session state, and mitigates locally without rerouting anything.
Hybrid runs the appliance always-on and keeps an upstream tier — either your ISP’s scrubbing service or a cloud provider — in reserve for attacks larger than your circuit can carry.
Note what the diagram makes obvious and marketing material tends not to: in the cloud-only architecture, all of your traffic — legitimate user sessions included, not just attack traffic — is inspected outside your jurisdiction, every day, whether or not you are under attack. That is not a criticism of cloud scrubbing; it is a structural property of it, and it is the property regulators ask about.
Latency: what “always-on” actually buys
The most common misconception in this decision is that cloud scrubbing is faster because the provider is bigger. Size determines capacity, not speed. What determines speed is whether the mitigation path is already carrying your traffic when the attack starts.
An on-demand diversion has to complete four steps before a single malicious packet is dropped: detect the attack, decide to divert, announce the change, and wait for the internet’s routing table to converge on the new path. The first two can be automated down to seconds. The last one cannot be automated away at all — BGP convergence takes as long as it takes.
Two honest qualifications. First, always-on cloud scrubbing removes most of this gap — if your traffic already flows through the provider, there is nothing to divert. Second, that mode changes the economics completely: you are now paying for continuous clean bandwidth rather than for occasional protection, and you have accepted the permanent cross-border inspection described above. The “fast” cloud option and the “cheap” cloud option are not the same option.
In steady state, the comparison inverts. An always-on cloud tier adds a round trip to the nearest scrubbing centre to every request — single-digit milliseconds if a PoP is in your city, 40–120 ms if the nearest one is on another continent. For a content site this is invisible. For a trading platform, a real-time bidding exchange, a telecom signalling path or an interactive game, it is the difference between a viable architecture and a non-starter.
Capacity: the hard ceiling nobody can engineer around
Here is the constraint that decides more architectures than any other, and the one that appliance vendors have the strongest incentive to soften: an on-premise appliance cannot filter traffic that has already saturated the circuit delivering it.
If you have a 10 Gbps access circuit and 25 Gbps arrives, the circuit is full before the appliance is consulted. It does not matter whether the appliance is rated for 20 Gbps or 200 Gbps. The bottleneck is upstream of it.
This is why “on-premise only” is a defensible architecture for a narrow set of organisations and an incomplete one for everybody else. It is defensible when your exposure is genuinely not volumetric — an internal service, a network segment behind an already protected perimeter, an environment where the threat model is application-layer abuse rather than bandwidth exhaustion. It is incomplete the moment a public-facing service can attract a booter service with a few hundred gigabits available for the price of a subscription.
The honest framing for a buyer is: an appliance protects the space below your circuit line, and everything you do above that line has to be bought from someone with a bigger pipe. Any vendor conversation that avoids this line should be treated as a warning sign.
Detection: what each tier can and cannot see
The two tiers fail in opposite directions, and understanding why explains most of the hybrid case.
A cloud scrubbing tier sees enormous breadth. It observes attack campaigns across thousands of customers, so a reflection vector first used against a gaming host in one region can be recognised the same week when it appears against a bank in another. What it does not have is depth about you. It does not know that your checkout flow normally sees 40 requests per session, that your API clients retry three times with a specific backoff, or that a burst of authentication attempts from one ASN at 03:00 is normal because that is when your partner runs a batch job.
An on-premise appliance is the mirror image. It sees every packet at line rate, holds full session state, and builds a behavioural baseline of your actual application. What it lacks is the outside world — it learns from your traffic alone.
This asymmetry is the strongest technical argument for hybrid, and it is a different argument from the capacity one. Even an organisation whose circuit could never be saturated gains from having both a global view and a local one. It is also the reason the two tiers should be tuned differently rather than identically: the upstream tier belongs on conservative, volumetric thresholds, while the fine-grained decisions belong to the layer that understands the application.
Sovereignty: whose law governs the inspection
A scrubbing tier cannot do its job without processing personal data. To distinguish a legitimate user from an attacker it must examine source IP addresses, request headers, user agents and, in most application-layer scenarios, session cookies or tokens. Where the provider terminates TLS to inspect Layer 7 — which is required for meaningful HTTP flood protection — it processes request bodies as well.
Under the GDPR, Türkiye’s KVKK, Saudi Arabia’s PDPL and the data-localization statutes of several Central Asian jurisdictions, all of that is personal data, and moving it to infrastructure in another country is a cross-border transfer requiring a lawful basis. This is not an argument that cloud scrubbing is illegal — with the right contractual instruments and transfer mechanisms it is routinely lawful. It is an argument that it is a question you must be able to answer, in writing, before a regulator or a customer asks.
The three questions that belong in every RFP:
- In which countries are the scrubbing centres that will handle our traffic? Not the provider’s global list — the specific PoPs your prefixes will be steered to.
- Is TLS terminated, and if so, where and by whom? A provider that inspects application-layer traffic without terminating TLS is limited to metadata; one that terminates it holds your users’ request contents.
- What is retained, for how long, and in which jurisdiction? Attack telemetry, sampled payloads and logs each have different answers, and vendors rarely volunteer them in the same document.
An on-premise tier does not answer these questions better. It removes them: the inspection happens on hardware you own, inside a facility you control, under one legal regime. For regulated buyers in the markets this site covers, that is frequently the deciding factor — and it is why “sovereignty” appears in tenders that never mention latency or packet rates. For an organisation already routing its prefixes through a network-layer cloud service, this is a matter of which tier carries everyday traffic by default rather than of replacing the upstream contract.
Cost: the five-year shape, not the first invoice
The three architectures have genuinely different cost shapes, and comparing them on year-one price systematically misleads.
Cloud scrubbing is opex that scales with two variables: the clean bandwidth you push through the provider, and — in many contracts — attack volume or event count. Its year-one number is almost always the lowest of the three, and its five-year number is the least predictable, because it moves with your growth and with what attackers decide to send you. Read the overage terms before the headline price; the gap between them is where cloud protection budgets break.
An appliance is capex plus an annual support renewal, and the renewal is a percentage of list that you should require in writing rather than assume, because it compounds across the whole term. Its five-year number is high but knowable on day one, which is exactly what a regulated procurement process wants. Its risk is on the other side: buy for today’s circuit and you re-buy when you upgrade, so size against the circuit you expect to have in year three, and check whether capacity is licensed rather than physical — a licence upgrade is a purchase order, a hardware upgrade is a project.
Hybrid costs less than the sum of its parts only if the cloud tier is sized for peaks rather than for steady-state traffic. The costing error to avoid is buying always-on cloud protection for your full clean bandwidth and an appliance: you then pay twice for the same everyday traffic and gain nothing over the cloud-only design.
| Appliance | Cloud scrubbing | On-premise appliance | Hybrid |
|---|---|---|---|
| Capacity ceiling | Provider backbone — effectively terabits | Your access circuit — nothing above it can be filtered | Circuit for everyday traffic, provider backbone for floods |
| Time to mitigation | Seconds always-on; minutes for on-demand diversion | Seconds from attack start, no rerouting | Seconds locally; upstream engages only when needed |
| Layer 7 depth | Good, but only where the provider terminates TLS | Strong — full session state and application baseline | Strong locally, volumetric handled upstream |
| Latency in steady state | Added round trip to the nearest scrubbing centre | None beyond appliance processing | None until diversion |
| Traffic jurisdiction | Inspected wherever the provider's PoPs are | Never leaves your premises | In-country by default, crosses only during diversion |
| Cost shape | Opex, scales with clean bandwidth and attack size | Capex plus support renewal, fixed and predictable | Both, with the cloud tier sized for peaks only |
| Operational load | Lowest — the provider tunes it | Highest — you own thresholds and policy | Moderate, plus diversion runbook and drills |
Every row is a generalisation across a product class; specific products differ. Always validate the two rows that decide your case — capacity ceiling and traffic jurisdiction — against the actual contract, not the datasheet.
When each architecture is the right answer
Cloud-only is right when your exposure is dominated by volumetric attacks, your traffic carries no regulated personal data or transfers are already covered, your latency budget is generous, and you have no team to operate an appliance. A small e-commerce business, a marketing site, a startup without a network engineer: cloud-only is not a compromise for these buyers, it is the correct answer.
On-premise-only is right when the traffic must not leave your jurisdiction, when latency is a product requirement, when your real threat is application-layer rather than volumetric, or when the environment is not internet-facing in a way that attracts terabit floods. Government systems, internal banking infrastructure, industrial networks and segments behind an already protected perimeter fall here.
Hybrid is right for nearly everyone else, and the specific trigger is simple: if a plausible attack against you could exceed your access circuit, and you cannot tolerate minutes of downtime while a diversion converges, you need both tiers.
Designing a hybrid that actually works
Most hybrid architectures that fail in production do not fail on capacity. They fail because the handover between the tiers was never designed, only assumed. Four things make the difference.
Three diversion triggers, not one. An automatic threshold on the appliance side (circuit saturation percentage, or a specific vector signature); an independent automatic threshold on the upstream side, because when the circuit fills your own signalling channel may be affected too; and a manual trigger your operations team can always exercise.
An out-of-band path for that manual trigger. If the only way to call your provider runs over the circuit under attack, you do not have an escalation procedure — you have a hope. A separate connection or a pre-agreed emergency channel must exist before the incident.
A documented fail-back. Diverting is the easy half. Returning traffic to the normal path without a second outage requires defined conditions, a defined owner and a tested sequence. Most organisations that have run a real incident discovered this during it.
Asymmetric tuning. Set the upstream tier conservatively — it should catch volume, not subtlety. Give the fine-grained decisions to the appliance, which knows what your application looks like on a normal Tuesday. Two tiers tuned identically produce twice the false positives with none of the additional coverage.
Procurement checklist
- Size against your circuit, not your peak traffic. The circuit is the line that decides which tier handles what.
- Ask for the five-year number in writing. Year 1 through year 5, including support renewal and overage terms.
- Name the scrubbing PoPs. Countries, not regions. Get it in the contract if jurisdiction matters to you.
- Test diversion end to end, including fail-back. Make the timing an acceptance criterion, not a demo.
- Validate Layer 7 with your own traffic. A replay of your real application beats any synthetic flood in a vendor lab.
- Check whether capacity is a licence or a chassis. It determines whether your next upgrade is a purchase order or a project.
Sources and further reading
We do not paraphrase paywalled analyst research or attribute figures to it. The list below is what to read, and what to ask a vendor to produce, rather than a set of claims made on someone else’s authority.
Analyst coverage of the category. Gartner covers this market in its Market Guide for DDoS Mitigation Solutions; Forrester has published The Forrester Wave: DDoS Mitigation Solutions; IDC publishes an IDC MarketScape for the segment. When a vendor claims analyst recognition, ask for the current edition rather than a screenshot — positions move between editions, and a two-year-old graphic is a marketing asset, not evidence. Note also that Gartner restricts public quotation of its research: a vendor quoting it on a public web page should hold reprint rights, and it is reasonable to ask to see them.
Standards and public technical guidance. NIST SP 800-189 on resilient interdomain routing; IETF RFC 8955 and RFC 8956 for BGP FlowSpec; RFC 9132 and RFC 8811 for DOTS, the standardised signalling protocol for cross-organisation mitigation requests — the relevant standard when you want the interface between your appliance and an upstream tier to be open rather than proprietary; ENISA’s Threat Landscape for European regulatory context.
Public incident and volume data. The quarterly DDoS reports published by Cloudflare and Akamai are the most widely cited open sources for attack volumes and vector mix; both are compiled from their own customer bases, which is worth remembering when reading their distributions. CVE-2023-44487 (HTTP/2 Rapid Reset) remains the clearest public example of a protocol-level flaw that no architecture choice could have isolated.
Regulatory texts. GDPR Chapter V on international transfers; Türkiye’s KVKK Article 9; Saudi Arabia’s PDPL implementing regulations on cross-border transfer; NIS2 Article 21 on risk management measures.
For buyers whose deciding constraint is the jurisdiction row rather than the capacity row, the practical shape is an always-on inline appliance handling everything below the circuit line — a device that carries layers 3 through 7 itself — paired with an upstream tier that is contractually engaged only above it. That keeps everyday traffic, and everyday personal data, inside one legal regime while still having an answer for the rare volumetric event.
Frequently asked questions
- Is cloud scrubbing always faster to deploy than an appliance?
- To deploy, yes — a DNS or BGP change against an appliance procurement cycle is not a fair fight. To mitigate, not necessarily. An always-on inline appliance starts mitigating within seconds of attack onset with no rerouting, whereas on-demand cloud diversion has to detect, decide, announce and wait for BGP convergence first. Always-on cloud modes remove most of that gap, but they also remove the "only pay when attacked" economics that made on-demand attractive.
- Can an on-premise appliance handle a terabit attack?
- No, and no vendor should claim otherwise. If the attack exceeds your access circuit, the circuit saturates before the appliance sees the traffic. Appliance throughput is irrelevant above that line. This is the single hard constraint that makes hybrid the default answer for anyone whose exposure includes volumetric attacks.
- Does cloud scrubbing create a data protection problem?
- It creates a data protection question that must be answered, not automatically a violation. A scrubbing tier cannot classify traffic without processing source IP addresses, headers and often session identifiers — all personal data under GDPR, KVKK and PDPL. If those PoPs sit outside your jurisdiction, you need a lawful transfer basis, a processor agreement and, in some sectors, an explicit regulatory position. On-premise mitigation removes the question rather than answering it.
- How much does hybrid cost compared with a single tier?
- Less than the sum of both, if the cloud tier is sized for peaks rather than for steady-state clean bandwidth. The common costing mistake is buying always-on cloud protection for full traffic volume while also running an appliance; the correct shape is an appliance sized to the circuit plus an upstream tier priced for the rare terabit event.
- Do I need two vendors, or can hybrid come from one?
- Hybrid works with one vendor and is easier to operate that way. What a single vendor cannot give you is failure independence: shared code, shared detection logic and a shared management plane mean both tiers can fail for the same reason. Whether that matters depends on your outage cost.
- What should a proof of concept actually test?
- Replay of your own application traffic, not synthetic floods; a full diversion and fail-back cycle with the upstream tier, timed end to end; behaviour under a state-exhaustion attack while the circuit is healthy; and the false-positive rate during a normal business peak. Anything a vendor demonstrates on their own traffic tells you about their lab, not your network.
- If we go hybrid, what should each tier be held responsible for in the contract?
- Split it at the circuit line and write the split down. Above it, only the upstream tier can help, so what you are buying there is capacity, a diversion trigger and a return path with times attached — an anycast service such as Cloudflare Magic Transit absorbs whole prefixes at a scale nothing in your rack can match. Below it — state exhaustion, application-layer floods, everything that never fills the pipe — the inline appliance owns the outcome, which is why L3–L7 depth in the same device matters more than a headline throughput figure; HARPP DDoS Mitigator sits in that lower half and changes nothing about the arithmetic above the line. Specify the two halves separately, then test the handover between them rather than each tier alone.
Published: August 2026
This guide is updated as vendors release new models and pricing. How we compare vendors