Architecture decision
Cloudflare Magic Transit On-Premise Alternatives: Keeping Traffic In-Country
Last updated: August 2026 · Cloud-first to in-country, honestly · Reading time ~18 min

Cloudflare Magic Transit protects whole IP subnets by attracting your prefixes into a large anycast network, filtering there, and returning clean traffic over a tunnel or interconnect. Its capacity and onboarding speed are genuinely hard to match on premises. The three defensible reasons to add an on-premise tier are unrelated to quality: which jurisdiction inspects your everyday traffic, the steady-state latency of an on-ramp/off-ramp path, and opex that scales with clean bandwidth versus fixed capex. An appliance answers all three below your access circuit and none of them above it — so the honest outcome is almost always an on-premise-primary hybrid, not a migration. The tier you add below the circuit has to answer all three by itself, which points at a single device covering L3 to L7 whose detection needs nothing from its manufacturer at runtime.
Most enquiries about on-premise alternatives to Cloudflare Magic Transit do not begin with a complaint. They begin with a change of circumstance: a new data protection position, a latency-sensitive service that did not exist at the last renewal, a finance department that wants a five-year number it can defend, or a regulator asking a question the current architecture cannot answer in writing. The service is working. Something around it moved.
That distinction matters, because it changes what a sensible evaluation looks like. If you are unhappy with a product, you compare products. If your constraints have changed, you compare architectures — and the honest conclusion of that comparison is usually not a migration at all. It is a rearrangement, in which everyday traffic comes back in country and the upstream tier stays contracted for the one thing no appliance can do.
This guide sets out what Magic Transit is genuinely good at, the three constraints that legitimately push work back on premises, and the shape of the resulting hybrid. The wider architectural argument — cloud versus on-premise versus hybrid on capacity, latency, cost and sovereignty — is developed in our architecture comparison; this article assumes it rather than repeating it.
What Magic Transit is good at, stated plainly
An evaluation that opens by understating the incumbent produces a bad decision, so it is worth being precise about what the service does well.
Magic Transit protects whole IP subnets rather than individual hostnames. Your prefixes are announced from Cloudflare’s network with BGP, traffic destined for them is attracted into that network and filtered there, and clean traffic is returned to your origin over a tunnel or a direct interconnect. That is a materially different product from a reverse proxy in front of a website: it covers protocols that are not HTTP, hosts that are not web servers, and infrastructure that has no hostname at all.
Three properties follow from that design, and they are real advantages rather than marketing.
Absorption capacity that no single site can hold. Because the network is anycast, a flood aimed at your prefix is not delivered to one location — it is pulled apart by internet routing and lands across many locations at once, each of which sees only a fraction of the total. This is the architectural reason large anycast networks are hard to saturate, and it is not something an organisation can reproduce inside its own facility at any price. Against the largest volumetric events, capacity close to the sources is the only mechanism that works, and there is no clever substitute for it.
Uniform mitigation everywhere. The filtering capability is a property of the network as a whole rather than of a particular scrubbing site, so there is no “nearest scrubbing centre is busy” failure mode of the sort that older centralised designs had to plan around.
Onboarding measured against a change window, not a procurement cycle. A BGP session and a tunnel can be stood up in days. There is no hardware lead time, no rack space, no power budget, no spares strategy and no shipping. For an organisation under attack this week, that is decisive, and it is the reason cloud-first became the default posture in the first place.
None of the three reasons discussed below contradicts any of this. They are questions about where the work happens and how it is paid for, not about how well it is done.
| Appliance | Cloud network-layer tier (Magic Transit shape) | On-premise appliance | On-premise-primary hybrid |
|---|---|---|---|
| Capacity above your access circuit | The reason to buy it — absorption far beyond any single circuit | None. A flood that fills the circuit is already past the appliance | Preserved, because the cloud tier is still contracted for exactly this |
| Where everyday packets are inspected | Wherever anycast routing lands them, subject to any localisation controls in scope | On hardware you own, in a facility you control, under one legal regime | In country by default; abroad only during a defined, logged diversion |
| Steady-state added latency | The on-ramp/off-ramp path, which is small with nearby ingress and material without it | Appliance forwarding only; no change of path | None until diversion; the cloud path cost is paid only while diverted |
| Time to mitigate at attack onset | Immediate while always-on; on-demand engagement adds detection and BGP convergence | Seconds, because the mitigation path already carries the traffic | Seconds locally; upstream convergence applies only to the volumetric case |
| Cost shape | Opex, moving with committed clean bandwidth and contract term | Capex plus support renewal — high but knowable on day one | Both, with the upstream tier scoped to peaks rather than to everyday traffic |
| Operational load on your team | Lowest — the provider tunes and operates the filtering | Highest — thresholds, baselines, policy and lifecycle are yours | Moderate, plus a diversion runbook that has to be drilled |
| Onboarding time | Fast — a BGP session and a tunnel, with no hardware lead time | A procurement cycle, a rack, a maintenance window and a tuning period | The appliance timeline, with the cloud tier already in place during it |
Rows describe architectural shapes, not the terms of any particular agreement. Capacity commitments, localisation options and pricing structures differ by contract and by market; the only authoritative source for your case is your own current quotation and the text of the agreement, not a general description of a product class.
The question an on-premise tier actually answers
Every mitigation architecture answers one structural question: on which side of a boundary is your traffic inspected, on an ordinary Tuesday when nothing is attacking you?
In a cloud-first posture, the answer is: outside, permanently, for all traffic. That is not a defect. It is the mechanism — a network cannot filter what it does not receive, and it cannot receive your traffic without your traffic going there. But it does mean that three otherwise unrelated concerns all resolve to the same architectural lever, and once one of them becomes binding, an on-premise tier stops being a preference and becomes the only instrument that moves the constraint.
Reason one: which jurisdiction inspects everyday traffic
This is the reason that appears most often in tenders, and it is the one most often argued badly by both sides.
The badly argued version is “cloud protection is illegal under our data protection law.” It generally is not. With an appropriate transfer mechanism, a processor agreement and honest disclosure, cross-border processing of traffic data is routinely lawful in most of the jurisdictions this site covers.
The well argued version is narrower and much harder to dismiss: filtering requires processing, processing abroad is a transfer, and a transfer is something you must be able to document, justify and disclose — permanently, for all users, not only during incidents.
Two nuances are worth stating precisely, because they are where evaluations usually go wrong.
Anycast means routing chooses the location, not your contract. The whole point of an anycast network is that a packet is delivered to whichever location the internet’s routing system considers nearest at that moment. That is what makes the absorption property work. It also means the ingress location for a given user population is a routing outcome rather than a term you negotiate. Providers offer localisation controls that constrain where processing occurs, and Cloudflare publishes such controls; what an evaluation must establish in writing is which controls apply to the network-layer service you are buying, as against the HTTP-layer services they were principally designed around, and what they do and do not constrain about first ingress. This is a genuine engineering distinction, not a gap — but it is one you should have documented rather than assumed.
“We do not store it” does not close the question. Under the GDPR, KVKK and PDPL, processing includes collection, examination, classification and transmission. Reading a packet in order to decide whether it is hostile is processing in its own right. Zero retention narrows the exposure; it does not remove the transfer, and it does not answer the question a supervisory authority actually asks, which is where the operation took place.
An on-premise tier does not answer these questions more convincingly. It removes them: one legal person, one facility, one legal regime, no processor chain, no sub-processor list, no additional recipient category to disclose. For public sector bodies, critical infrastructure operators and regulated financial institutions, that difference is frequently the whole decision — and it is why “sovereignty” appears in tender documents that never mention packet rates.
The trap to avoid is the mirror image of the first bad argument. If you keep an upstream tier for volumetric events — and you should — then “no data ever leaves the country” is a false statement in your privacy notice. The correct disclosure describes an exceptional, conditional, logged transfer with a defined trigger. That is both accurate and easy to defend.
Reason two: steady-state latency and the shape of the path
A network-layer cloud service is an on-ramp/off-ramp architecture. Traffic enters the provider’s network, is filtered, and is delivered to your origin over a tunnel or an interconnect. The latency consequence of that is entirely a question of geography, and both of the following are true depending on where you sit.
Where the ingress point for your users is in the same metropolitan area as your origin, the added path is short and most applications will never notice it. Where it is not — where a domestic user population’s traffic is drawn to a point outside the country and then returned over a tunnel to a server that was a few milliseconds away to begin with — the addition is structural and permanent. It is paid on every request, every day, whether or not anyone is attacking you.
For a content site, a corporate portal or a document workflow, this is invisible and irrelevant. For a small and growing set of services it is decisive: exchange and trading platforms, real-time bidding, interactive multiplayer, telecom signalling paths, remote industrial control, and increasingly interactive AI inference endpoints where users are sensitive to time-to-first-token. In those environments the latency budget is a product requirement, and an architecture that spends part of it permanently on a security function is competing against one that spends none.
Two operational details belong in the same conversation, because they surprise people after signature rather than before.
Encapsulation costs payload. Traffic returned over a GRE tunnel carries additional headers, so the space available for application data in each packet shrinks. Without correct MSS clamping and path MTU handling, the symptom is not a clean failure — it is a minority of connections that stall on large transfers while everything else looks healthy. This is well-documented and entirely manageable, but it is configuration work you own, and it is the single most common cause of “the tunnel is up but something is odd” tickets.
Asymmetric paths complicate diagnosis. In an on-ramp/off-ramp design the inbound and outbound paths are not mirror images, which is fine until you are trying to reason about a performance complaint under time pressure. Budget for the observability to see both halves.
An on-premise appliance changes none of the path. It sits inline between your border router and your services, and the only latency it adds is its own forwarding time. For a latency-bound service, that is the entire argument.
Reason three: opex that scales with clean bandwidth against fixed capex
The third reason is the one finance raises, and it is not really about which option is cheaper. It is about which shape of number the organisation can plan against.
A cloud tier is opex. Its cost moves with the clean bandwidth you commit to and with the commercial terms of your agreement. That is an excellent fit for an organisation whose traffic is unpredictable, whose growth is uncertain, or which prefers to hold no assets. It is a poor fit for one whose budget cycle demands a five-year figure approved in advance, and it becomes uncomfortable when traffic grows steadily and predictably — because then you are paying an increasing amount, indefinitely, for the everyday inspection of traffic that never had an attack in it.
An appliance inverts both properties. It is capital expenditure plus a support renewal: higher over five years in many cases, but knowable on the day you sign, insensitive to traffic growth up to the capacity you bought, and structured the way regulated procurement processes prefer. Its risk sits elsewhere — buy for today’s circuit and you re-buy when you upgrade, so size against the circuit you expect in year three, and establish whether capacity is licensed or physical. A licence uplift is a purchase order; a chassis change is a project.
The point the diagram makes is the one most business cases miss. The appliance is not evaluated against the cloud subscription alone. It is evaluated against the subscription plus the downstream consequences of everyday attack traffic reaching equipment that is licensed by session count, inspection throughput or events ingested. Whether those savings are large or trivial depends entirely on how often you are actually attacked and how your existing licences are structured — which is why this model is only meaningful with your own numbers substituted into it. Our total cost of ownership guide works through the arithmetic.
What an appliance cannot do, and why this is not a migration
Here is the constraint that should end any conversation framed as full replacement.
An on-premise appliance cannot filter traffic that has already saturated the circuit delivering it. If you hold a 10 Gbps access circuit and 40 Gbps arrives, the circuit is full before the appliance is consulted. Whether the appliance is rated for 20 Gbps or 200 Gbps is irrelevant above that line, because the bottleneck is upstream of the device.
This is why the honest recommendation for an organisation coming from a cloud-first posture is not “replace Magic Transit with an appliance.” It is: change which tier is the default, and keep the other one for the case it uniquely solves.
Concretely, the destination architecture is an always-on inline appliance handling everything below the circuit line — all everyday traffic, all application-layer abuse, all state-exhaustion attempts, all the protocol-level nuisance that never approaches your bandwidth limit — paired with an upstream tier contractually engaged above it. Appliances that consolidate layers 3 to 7 into one unit fit this position directly, because it is the everyday inspection that has to stay in country and the everyday inspection that has to add no path.
What changes commercially is not usually the existence of the upstream contract but its scope. A tier sized for peak events is a different commitment from a tier carrying your steady-state clean bandwidth, and that difference is where the cost predictability argument actually lands.
Designing the on-premise-primary hybrid
Reversing the default is a design exercise, not a configuration change. Four things decide whether it works in production.
Diversion has to have more than one trigger. An automatic threshold on the appliance side — circuit utilisation, or a specific vector signature. An independent automatic threshold upstream, because when your circuit fills, your own signalling path may be impaired too. And a manual trigger your operations staff can always exercise. One trigger is a single point of failure wearing a runbook.
The manual trigger needs an out-of-band path. If the only way to reach your provider runs across the circuit under attack, you do not have an escalation procedure. Agree the channel, the authorised callers and the authentication method before the incident, not during it.
Fail-back must be documented and drilled. Diverting is the easy half. Returning traffic to the normal path without causing a second outage needs defined conditions, a named owner and a tested sequence. Organisations that have run one real incident invariably discovered this during it.
Tune the tiers asymmetrically. The upstream tier should be conservative and volumetric — it is there to catch volume, not subtlety. Fine-grained decisions belong to the layer that knows what your application looks like on a normal Tuesday. Two tiers tuned identically produce twice the false positives and no additional coverage.
One further design decision is worth making deliberately: whether the interface between your appliance and the upstream tier is proprietary or standards-based. DOTS — specified in RFC 9132 with the architecture in RFC 8811 — exists precisely so that a mitigation request can cross an organisational boundary without both ends coming from the same manufacturer. Where an upstream provider supports it, it keeps your future options open at no cost to the present design. Where it does not, at least make the signalling method an explicit contract term rather than an assumption.
Practical migration sequence
The sequence below keeps you protected throughout, which matters because the riskiest moment in any such project is the one where neither tier is clearly in charge.
- Measure the current path first. Steady-state latency from each significant user population, current clean bandwidth, and the real attack history against your prefixes. Without a baseline you cannot demonstrate improvement, and you cannot size the appliance.
- Establish the announceable address blocks. BGP-based diversion needs prefixes that are actually routable — on the public internet, that means /24 or shorter for IPv4. If your addressing does not currently support that, fix it before designing anything else.
- Install the appliance inline in monitoring mode. Let it learn a behavioural baseline across at least one full business cycle, including your seasonal peak if one is approaching, while the cloud tier remains exactly as it is.
- Enable local mitigation while the cloud tier is still always-on. Nothing is at stake yet; you are validating that the appliance’s decisions match what you would have made.
- Change the default. Move the upstream tier to an on-demand posture, with everyday traffic taking the direct path. This is the step that delivers the residency and latency outcomes, and it is the step to schedule in a quiet window.
- Drill the diversion, then drill the fail-back. Timed, with named owners, and repeated at a defined interval thereafter. A diversion path that has never been exercised is a plan, not a capability.
- Update the privacy notice and processing records. The transfer is now conditional rather than continuous. Say so accurately — including the trigger conditions — rather than claiming it no longer occurs.
When staying cloud-first is the right answer
A comparison that always concludes “change” is not a comparison. Several situations point clearly the other way.
Your exposure is overwhelmingly volumetric and your circuit is modest. If the attacks that actually reach you routinely exceed what your access circuit can carry, the tier that matters is upstream, and an appliance would spend most of its life idle behind a saturated link.
Your traffic carries no regulated personal data, or the transfer question is already answered. If the residency argument does not bind, it should not be manufactured. One fewer question is worth something; it is not worth a capital project on its own.
You have no team to operate an appliance. Thresholds, baselines, false-positive management, firmware lifecycle and out-of-hours incident response are real recurring work. An unmaintained inline device is worse than no inline device, because it is also a failure domain.
Your latency budget is generous and your traffic is unpredictable. The two properties that make opex uncomfortable — steady growth and the need for a fixed five-year figure — are exactly the two that may not apply to you.
The decision framework reduces to a single test. Ask which of the three constraints is actually binding in your organisation right now: the jurisdiction of everyday inspection, the steady-state latency of the path, or the shape of the cost line. If none of them is, stay where you are and revisit at renewal. If one of them is, the instrument that moves it is an on-premise tier — and the correct scope of that project is the traffic below your circuit line, not the traffic above it.
Sources and further reading
We do not paraphrase paywalled analyst research or attribute figures to it, and we do not restate vendor capacity or pricing claims. The list below is what to read, and what to ask a vendor to produce in writing.
Vendor documentation, read directly. Cloudflare’s own developer documentation is the authoritative description of how Magic Transit attracts prefixes, what transport options exist for returning clean traffic, and which data localisation controls are available and to which services they apply. Read the current documentation rather than a summary — including this one — and get the answers that matter to your case restated in your agreement. Cloudflare’s quarterly DDoS threat reports remain among the most widely cited open sources for attack volumes and vector mix; like all such reports they are compiled from the publisher’s own customer base, which is worth holding in mind when reading their distributions.
Analyst coverage of the category. Gartner covers this market in its Market Guide for DDoS Mitigation Solutions; Forrester has published The Forrester Wave: DDoS Mitigation Solutions; IDC publishes an IDC MarketScape for the segment. Ask any vendor claiming analyst recognition for the current edition rather than a screenshot — positions move between editions. Note also that Gartner restricts public quotation of its research: a vendor quoting it on a public web page should hold reprint rights, and it is reasonable to ask to see them.
Standards and public technical guidance. RFC 9132 and RFC 8811 for DOTS, the standardised signalling protocol for cross-organisation mitigation requests — the relevant standard when you want the interface between your appliance and an upstream tier to be open rather than proprietary. RFC 8955 and RFC 8956 for BGP FlowSpec. RFC 7454 (BCP 194) on BGP operations and security, and RFC 4786 (BCP 126) on the operation of anycast services, which together explain why routing rather than contract determines where an anycast packet lands. RFC 2784 and RFC 2890 for GRE, and RFC 4459 on MTU and fragmentation issues with in-the-network tunnelling — the document behind every MSS clamping recommendation you will be given. NIST SP 800-189 on resilient interdomain routing, and the MANRS actions for routing security hygiene on your own announcements.
Regulatory texts. GDPR Chapter V on international transfers; Türkiye’s KVKK Article 9; Saudi Arabia’s PDPL implementing regulations on cross-border transfer; NIS2 Article 21 on risk management measures. Where your sector has its own supervisor — banking, telecoms, energy — its outsourcing and information systems rules will usually have something specific to say about foreign elements in a security service, and that text governs over any general analysis.
Related reading on this site. The architecture comparison develops the cloud, on-premise and hybrid trade-off in full; the total cost of ownership guide covers the cost-line model referenced above; and the multi-vendor architecture guide addresses whether the two tiers should come from the same manufacturer.
Frequently asked questions
- Is there an on-premise product that replaces Magic Transit?
- Not in the sense buyers usually mean. An appliance can replace the everyday inspection function completely, and for most organisations that is the majority of the value they consume. What it cannot replace is absorption capacity above your access circuit: if a flood larger than your circuit arrives, the circuit saturates before the appliance is consulted, and no appliance rating changes that. Anything marketed as a full replacement for a large anycast network is describing a different physical situation from the one you are in.
- Does moving to an appliance mean leaving Cloudflare entirely?
- It should not, and treating it as one decision is the most common planning error. Network-layer transit protection, HTTP-layer services for your web properties, DNS and authoritative name service are separable purchases on separable renewal calendars. A great many organisations that add an on-premise tier keep an upstream contract specifically for the volumetric case, and keep their web-facing services exactly where they were. Establish which layer you are actually changing before asking anyone for a quotation.
- Why would data residency be a problem if traffic is only inspected, not stored?
- Because under the GDPR, KVKK, PDPL and comparable regimes, inspection is processing. Deciding whether a packet is hostile requires reading source IP addresses and, for application-layer defence, headers and session identifiers. A zero-retention policy reduces the risk profile; it does not remove the transfer. The question a regulator asks is where the processing happens, not how long the result is kept — and it is a question you should be able to answer in writing before it is asked.
- How much latency does a cloud network-layer tier actually add?
- It depends entirely on the geography of the path, and both extremes are real. Where your users' traffic enters the provider's network in the same metropolitan area as your origin, the addition is small enough that most applications will not notice it. Where the nearest ingress for a given user population is on another continent, or where the return path hairpins through a distant point, the addition is large enough to change an application's viability. Measure it on your own path with your own traffic; it is not a number anyone can quote generically.
- Does an on-demand diversion model give me the best of both?
- It gives you in-country everyday inspection with upstream capacity in reserve, which is exactly the point — but it is not free. Engaging an upstream tier on demand reintroduces detection, decision and BGP convergence time before the first malicious packet is dropped upstream. The appliance covers that window for anything below the circuit line, which is why the pairing works; but the diversion has to be triggered by more than one mechanism, drilled, and fail-back tested.
- Can I announce a prefix smaller than a /24 for on-demand diversion?
- Not on the public internet. Prefixes longer than /24 are widely filtered and are not reliably routable globally, which constrains any BGP-based diversion architecture regardless of provider. If your address space is not organised into announceable blocks, that is a network design task to complete before the mitigation design, not after.
- What should the proof of concept measure?
- Four things, all on your own traffic. Steady-state latency on the current path against the proposed one, measured from the user populations that matter. Behaviour under a state-exhaustion attack while the circuit is healthy, because that is the failure mode the perimeter firewall loses to. A full diversion and fail-back cycle with the upstream tier, timed end to end and treated as an acceptance criterion. And false positives during a genuine business peak, not a quiet afternoon.
- If we keep Magic Transit, what is the on-premise tier actually being bought to do?
- Not to replicate the anycast tier, and any evaluation that scores it that way will conclude it is redundant. It is bought for the traffic that never leaves your access circuit: everyday inspection that stays in your jurisdiction, application-layer classes that sit far below the threshold where diversion is worth triggering, and continuity of protection during the interval when the upstream tier is being engaged. Written that way the requirement is narrow: permanently in path, L3 through L7 in one unit, verdicts reached locally. That is a different specification from the one a scrubbing centre is built to, and it sorts the market quickly — Corero SmartWall is deliberately narrow and low-touch inline, HARPP DDoS Mitigator carries the full L3–L7 span in the one device. Neither will survive an attack larger than the circuit, which is precisely why the cloud tier stays.
Sources
- Magic Transit — DDoS protection for networks
Cloudflare · vendor documentation · accessed 2026-08-15
- SmartWall ONE — DDoS protection
Corero Network Security · vendor documentation · accessed 2026-08-15
Published: August 2026 · Last reviewed: August 2026
Reviewed means the sources above were re-read on that date; the text is only reissued when something material changed.
This guide is updated as vendors release new models and pricing. How we compare vendors