Skip to content

Procurement guide

DDoS Mitigation Buyer's Guide for ISPs and Telecom Operators

Last updated: August 2026 · Specifying and evaluating operator-grade mitigation · Reading time ~21 min

A wide operator ring with many thin customer lines radiating inward, the attack pressure bearing on the ring itself rather than on any single line: the specification is written against your own edge.

An operator's DDoS specification is written against your own edge, not your customers' circuits: peering and transit capacity, backhaul to the scrubbing cluster and packet rate set the ceiling. Asymmetry limits what stateful inspection can honestly claim, BGP and FlowSpec integration decide how fast you divert, and per-tenant policy and reporting decide whether you have a product or a cost centre.

There are two quite different documents an operator can write about DDoS mitigation. One is a build guide: how to place the scrubbing cluster, how to solve the return path, how to tune the detection thresholds. The other is a procurement document: what to specify, how to score the answers, and what you are allowed to promise the customer afterwards. This is the second kind. If your task this quarter is to construct the thing, our scrubbing centre build guide covers the engineering sequence. If your task is to write the tender, evaluate the bids and defend the resulting commitment to your commercial director, read on.

The two documents diverge on one point above all. The builder asks how do I make this work. The buyer has to ask what does this platform do at the boundaries, because the boundaries are where the SLA is tested and where the money is lost. Almost every unpleasant surprise in an operator DDoS programme comes from a limit that was known to the vendor, absent from the specification, and discovered during an incident.

What you are actually buying

An operator does not buy a DDoS product; it buys four separable things that vendors bundle under one name.

The first is detection — the ability to see, from flow telemetry or from inline inspection, that a specific customer prefix is under attack, and to say so within a time budget you can put in a contract. The second is traffic steering — the machinery that redirects the affected traffic to the place where it can be cleaned, and returns the clean traffic to the customer. The third is mitigation — the countermeasures themselves, their throughput, and their behaviour under the traffic conditions your network actually produces. The fourth, and the one most often left out of technical evaluations, is the tenancy and reporting layer that turns all of the above into a line item on an invoice.

A tender that scores only the third of these will select a fast box that you cannot sell. A tender that scores all four tends to select differently, because the vendors that are strongest on raw mitigation throughput are not always the ones that are strongest on per-tenant separation and per-tenant reporting.

Three delivery models, and the table to put in front of your board

Before the technical specification comes a commercial decision that changes what the specification needs to say. You can build and operate scrubbing capacity yourself, white-label capacity from a scrubbing partner, or resell what your own transit provider already sells. Each model puts the ceiling, the jurisdiction and the escalation path in a different place, and where the capital for the first is not available this year, the measures that cost engineering time rather than money come before any of the three.

ApplianceBuild your own capacityWhite-label a scrubbing partnerResell your transit provider's service
Time to first revenueLongest — design, procurement, integration and drills before the first invoiceMedium — commercial and portal integration dominateShortest — often a contract amendment and a price list
Ceiling on what you can protectYour peering and transit edge, your backhaul and your packet-rate budgetThe partner's capacity, minus whatever they have already soldThe provider's capacity, on terms you do not control
Control over per-tenant policyComplete — you write the baselines and the countermeasure orderPartial — usually a policy template menu rather than free configurationMinimal — you raise tickets like any other customer
Where customer traffic is inspectedWherever you place the cluster, under your own jurisdiction and contractsIn the partner's footprint, under the partner's jurisdictionIn the provider's footprint, under the provider's jurisdiction
Per-customer reportingYours to build or to buy as a product featureDepends on whether the partner offers white-label portals and an APIUsually absent or unbranded — the hardest gap to close commercially
Margin structureCapital and operating cost up front, margin improves with tenant countWholesale-to-retail spread, predictable but cappedThin resale margin, and the provider owns the customer relationship
Typical failure of the arrangementCapacity sold above the real edge, so physics breaks the commitmentPartner outage or capacity contention that you cannot diagnoseEscalation path runs through a third party during an incident

The three models are not mutually exclusive and most operators end up combining them: own capacity for the volumes they can absorb, and a contracted overflow path above that threshold. The decisive question is not cost per gigabit but which model lets you keep a per-customer report and an escalation path you actually control.

Most operators outside the largest markets end up in a hybrid position: their own capacity handles the attack volumes that are actually common against their customer base, and a contracted overflow path handles the rest. That is a defensible design. What is not defensible is leaving the threshold between the two undocumented, because the threshold is exactly what the customer is buying.

Capacity: specify against your edge, not your customers’ circuits

This is the most common sizing error in operator procurement, and it is expensive in both directions — you either buy too much or you sell more than you can deliver.

An enterprise sizes mitigation against its access circuit, because the circuit is the physical limit on what can reach it. An operator has no such convenient number. The figure that governs you is how many bits and how many packets can physically enter your network across the sum of your peering and transit ports. That is the true upper bound on the attack you might be asked to absorb, and it is usually far larger than any single customer’s circuit.

Where the on-premise layer stops being able to help Your 10 Gbps access circuit Circuit already saturated — upstream only On-premise appliance mitigates 2 Gbps 8 Gbps 25 Gbps 120 Gbps 1 Tbps+ Attack volume (log scale)
On the customer side the limit is the access circuit. On the operator side the limit is the aggregate of peering and transit ports — a much higher line, and the one your specification has to be written against.

Three further ceilings sit underneath that headline number, and a tender that ignores them will produce a platform that cannot deliver its own datasheet.

Backhaul to the scrubbing cluster. Diverted traffic has to travel from the edge routers to wherever the cleaning happens. Whatever the cluster’s nameplate rating, you cannot divert more than the backbone can carry to it. In many designs this is the real bottleneck, and it is the line most easily forgotten in the budget because it is a transport cost rather than a security cost.

Packet rate. Bit rate alone flatters every platform. Attacks composed of small packets hit the packet-processing limit well before they approach the quoted bit rate, so a datasheet expressed only in gigabits per second describes the easy case. Require the packet-per-second figure in writing, require it with the countermeasure set you intend to run enabled, and run acceptance testing at the smallest frame size the platform will realistically encounter.

Concurrency. Sizing for the largest single attack you can imagine is optimistic. Attacks arrive in campaigns, and it is entirely ordinary for the same campaign to target several of your customers within the same hour. Write down an explicit assumption about how many tenants may be under simultaneous mitigation, and size against that assumption rather than against a single event.

To make the arithmetic concrete — and this is a worked example only, which you must redo with your own telemetry — suppose an operator has an aggregate external edge of 200 Gbps across peering and transit, 2 × 100 Gbps of backhaul to a single scrubbing site, and a platform rated at a packet rate that your own testing shows to be reached at roughly 60 Gbps of 64-byte traffic. The protection level you can honestly advertise is derived from the smallest of those three constraints under the traffic mix you expect, not the largest, and if you commit to protecting two tenants concurrently, their combined figure has to stay under the same line. None of the numbers above are norms; they are placeholders for yours.

The commercial consequence of getting this wrong deserves to be stated plainly. If you sell a protection level above your real edge capacity, an attack of that size fills your transit ports before it ever reaches the scrubbing cluster. The damage is not confined to the customer under attack — every customer behind those ports is affected, you fall outside the traffic profile you committed to your own upstream, and where transit is billed on a 95th-percentile basis the attack traffic converts directly into your own cost. A published blackhole threshold is not a weakness in a proposal; it is the mark of an operator who has done the arithmetic.

Diversion signalling: what the BGP integration has to do

Diverting traffic is conceptually simple — announce a more specific prefix from the scrubbing site and the longest-prefix-match rule does the rest — and operationally full of sharp edges. The RFP should specify the mechanics, not just the outcome.

Session design. State how the mitigation platform is expected to inject routes: a dedicated eBGP session to your route reflectors or to specific edge routers, with session authentication, and with a strict inbound policy on your side. The controller is a device that can, by design, redirect production traffic, so the session it speaks on needs the same protections as any other external session — a maximum-prefix limit, a prefix-length filter, an allow-list of the prefixes it may ever announce, and community-based tagging so that diversion routes are identifiable everywhere in your network. RFC 7454 (BCP 194) is the right baseline document to reference in the tender for BGP operational security generally.

Withdrawal behaviour. Ask what happens to announced diversion routes if the controller loses management connectivity, if the process restarts, or if the operator forgets. A dead-man’s-switch behaviour — routes expire unless refreshed — is a materially different safety posture from routes that persist until explicitly removed. Both are defensible; only one of them matches your runbook, and you should know which before you sign.

Prefix length outside your AS. Inside your own AS you may announce a more specific prefix of any length, because the announcement never leaves. The moment diversion has to be signalled beyond your AS boundary, the widespread operational practice of filtering prefixes longer than /24 in IPv4 and /48 in IPv6 becomes binding. This is a convention rather than a standard, but it constrains you all the same, and for a hosting provider that assigns individual addresses to customers it means the address plan has to be designed around the unit of protection from the beginning.

RPKI and the announcement you need most. If a diversion announcement will be seen outside your AS, the maxLength value in the ROAs covering those prefixes must permit the longer prefix you intend to announce. If it does not, the announcement is invalid at every network performing origin validation and is discarded — precisely at the moment you need it. This sits in direct tension with the advice in RFC 9319 (BCP 185) to keep maxLength values tight, and the two requirements have to be planned together rather than by two teams separately. Put “the diversion announcement is accepted and visible externally” into the acceptance test plan.

Upstream signalling. What you can ask your own upstreams to do is limited to the community set they publish. Remotely triggered blackholing is the one primitive available almost everywhere: RFC 7999 defines the well-known BLACKHOLE community, the destination-based method is described in RFC 3882, and the source-based variant with unicast reverse path forwarding in RFC 5635. Be honest internally about what this is — blackholing is not protection, it is surrender of the target to protect everything else. Collect each upstream’s accepted community list, accepted prefix lengths and committed response times at contract stage, not during an incident.

Time from attack start to full mitigation On-demand cloud diversion Detection Decision / announcement BGP convergence Mitigating Always-on inline appliance Detect Mitigating 0 1 min 2 min 3 min 4 min 5 min Indicative ranges. Diversion time depends on the provider, the announcement method and the state of the routing table — always-on cloud modes are faster than the on-demand path shown here.
Time to mitigation in an out-of-path design is a sum: telemetry export, detection decision, diversion announcement and convergence. You cannot write a defensible SLA number until you know which of those components you can actually shorten.

Standards-based signalling with customers. If some of your customers run their own on-premises appliances, the request-for-mitigation signalling between them and you does not have to be proprietary. DOTS is the relevant work: RFC 8811 for the architecture, RFC 9132 for the signal channel and RFC 8783 for the data channel. Asking for it in the tender costs you nothing and preserves the option of a hybrid arrangement in which the customer’s device and your layer cooperate without either side owning the other’s interface.

Asymmetric routing and the limits of stateful inspection

This is the technical area where operator-grade evaluation diverges most sharply from enterprise evaluation, and where vendor datasheets are least useful.

In a network with multiple transit and peering points, the two directions of the same session routinely traverse different edge routers. Closest-exit routing makes this the normal case rather than an exception. Out-of-path diversion then makes asymmetry structural rather than incidental: during mitigation you divert only the ingress direction, while egress continues to follow the ordinary path. Even inside the scrubbing cluster, the way traffic is distributed across member devices can split what a single inspection engine sees.

The consequence is direct. An inspection engine that observes only one direction cannot genuinely hold connection state. It cannot confirm that a TCP handshake completed, cannot follow sequence-number progression, cannot see the session close. Faced with this, a device running a strict state machine does one of two things: it drops legitimate traffic, or it quietly disables state checking. Either behaviour may be acceptable, but you need to know in advance which one you have bought.

What works on a one-directional path is the family of countermeasures where the device itself generates the response and evaluates the client’s reaction: SYN cookie-style challenges, source validation based on retransmission behaviour, protocol conformance checks, and rate limiting. These do not require the return direction because the device supplies the return direction itself. Signature and pattern matching on ingress also survives asymmetry. What does not survive is anything that requires observing the server’s answer.

Turn that into specific tender questions:

  1. Which countermeasures require symmetric visibility? Ask for the list explicitly, not for a reassurance. Any capability on that list must be scoped in your service description to deployments where both directions are guaranteed to pass the same point.
  2. What is the default behaviour when the handshake is never observed? Drop, pass, or fall back to a challenge? The answer determines your false-positive exposure during every diversion.
  3. Is load distribution across cluster members per-flow? Per-packet distribution breaks state tracking and fragment handling on its own.
  4. How are fragments handled across cluster members? Non-initial fragments do not carry transport-layer port numbers, so a hash that includes ports will not send them to the same member as the first fragment. Ask how the platform reassembles or redirects them, and test it — fragmentation-based attacks exist precisely because this is hard.
  5. What is the state model, and can it be exhausted? A device that allocates per-flow state has a state table, and a state table is an attack surface in its own right. Ask what happens when it fills, and whether per-tenant limits can be applied to it.
  6. What does the service description say? Whichever features degrade under asymmetric traffic, the sales team needs to know before the contract, not the operations team afterwards.

Forcing symmetry is possible — pinning a customer to a single edge point, or running the egress direction through the scrubbing path as well — but it is a deliberate cost in capacity and latency. Treat it as a per-customer option with a price, not as a network-wide design default.

FlowSpec: sharp, and worth specifying carefully

BGP FlowSpec distributes n-tuple match rules over BGP: destination and source prefix, protocol, port ranges, packet length, fragment bits, TCP flags. On the action side it offers traffic-rate limiting (with a rate of zero acting as a discard), redirection to a VRF, and marking. RFC 8955 covers IPv4 and RFC 8956 IPv6.

For an operator the attraction is precise: a reflection attack can be cut at the edge router in hardware, without hauling the traffic to the scrubbing cluster at all. That protects the backhaul capacity discussed above, which is one of the most expensive lines in the capacity model.

The caution is equally concrete, and the tender should reflect it. A malformed rule propagates across the whole edge within seconds. Hardware filter resources on line cards are finite, and behaviour when rules exceed them varies by platform and vendor — unpredictability is itself the risk. The validation procedure in RFC 8955 ties a rule to the neighbour from which the best unicast route for the destination prefix was learned, which is correct for security but narrows what can be expressed in multi-provider environments. And implementation coverage differs: which combinations of match and action are supported, and which of them run in hardware rather than punting to the control plane, is a per-platform question.

Specification items worth writing down: support for RFC 8955 and RFC 8956; a documented list of match and action combinations that the vendor confirms run in hardware on your line-card inventory; a configurable hard cap on generated rule count; automatic expiry on every rule; the ability to scope generated rules per tenant; a full audit trail of what was generated, by whom or by what, and when; and a single, tested kill switch that withdraws everything. Operate FlowSpec inside your own AS first and treat any upstream acceptance as a bonus rather than a load-bearing element of the design.

Multi-tenancy: the criterion that turns cost into product

Build everything above and omit this, and you have a good network hygiene tool. You do not have something you can invoice. Three capabilities separate the two.

Per-tenant policy. Every customer’s normal is different. A tenant hosting game servers has a UDP profile that cannot be governed by the same thresholds as a bank whose traffic is overwhelmingly TCP on port 443. On shared hardware you need a separate learned baseline, a separate threshold set, a separate countermeasure order and a separate allow-list per customer — and you need policy templates, or fifty customers become fifty hand-written configurations that no one can maintain.

Isolation between tenants. An attack aimed at one tenant must not weaken another tenant’s protection by consuming shared resources. That means per-tenant limits at the level of the state table, the rule capacity and the processing budget. The question to ask a bidder is not “how many tenants does it support” — that number is marketing. The question is “how can one tenant affect another”, and the answer should be specific about which resources are partitioned and which are shared.

Per-tenant visibility. The customer must see only their own prefixes; must be able to follow the attack timeline, the vector breakdown and the volume dropped versus passed from their own portal; must be able to export the report; and must be able to open it to their own team under role-based access control. An API matters too, because technical customers want the data in their own monitoring system, and for them its absence is a reason to buy elsewhere.

The third item is the most commercially underrated. Your customer never sees your detection engine. They see the portal. A legible report showing that an attack was stopped is more persuasive at renewal than any claim about the engine’s sophistication, and the same report is what the customer hands to their own auditors and their own regulator.

The hardware criterion follows directly from this: separate protection profiles and separate reporting per tenant on shared hardware. Trying to sell a service on a product that does not meet it produces a configuration burden that grows faster than the customer base.

What can honestly be committed in an SLA

The value of an SLA lies as much in what it excludes as in what it promises. Commit to what you measure and control:

  • Time to diversion. From the detection event to the diversion announcement. Where this is automated it is a committable quantity.
  • Time to mitigation. From diversion to the mitigation policy being applied. Write into the contract where the measurement point is.
  • First response to a false-positive report. The false-positive rate cannot be committed; the number of minutes before an engineer looks at the policy can be. Define who is entitled to raise such a report and how their identity is verified.
  • Availability of the service itself. The availability of the scrubbing platform and the portal is a different thing from the availability of the customer’s service. Write them separately.

What cannot be committed should be written with equal clarity: that every attack will be stopped, that no legitimate request will ever be affected, that end-user experience will be unaffected, and that an attack larger than your edge capacity will be mitigated. Alongside these, define the mechanics of measurement — which clock starts when, whose telemetry is authoritative, which situations are out of scope (maintenance windows, customer-side misconfiguration, events above the published threshold) — and what remedy applies. In practice the remedy is a service credit capped at the monthly fee, and saying so up front is better than discovering the ambiguity in a dispute.

Supply, dependency and lifecycle criteria

A mitigation platform is a long-lived asset in a market where the vendor relationship can be interrupted for reasons that have nothing to do with the product. For operators outside the United States and Western Europe this is not a hypothetical procurement concern, and it belongs in the scoring matrix rather than in a footnote.

What survives if the vendor relationship is cut Keeps working Degrades over weeks Stops Hardware keeps forwarding packets Keeps working Locally trained detection keeps working Keeps working Cloud threat feed goes stale Degrades over weeks Licence renewal blocked — features expire Stops Support, RMA and spare parts stop Stops Firmware and security patches stop Stops The order matters more than the list. A product whose core detection depends on a vendor cloud moves the third row upward — it stops being a degradation and becomes an outage.
The order of degradation matters more than the list. Hardware keeps forwarding and locally trained detection keeps working; licences, patches, support and any cloud-delivered feed do not. A product whose core detection lives in a vendor cloud moves the third row upward, turning a slow degradation into an outage.

Concrete questions for the tender:

  • Does core detection function without reachability to a vendor-operated cloud? If the answer is no, you have bought a service dependency, not an appliance, and it should be priced and risk-assessed as one.
  • What does licence expiry disable? Mitigation itself, or only feature updates? Ask for the behaviour, in writing, not for the intention.
  • Export licensing and sanctions exposure. Which jurisdictions can interrupt supply, support or payment for this product, and what is the vendor’s documented position on continuing support in your market?
  • Spares, RMA and lead times in your region. Where is the spares depot, what is the replacement commitment, and does it survive a shipping disruption?
  • Support language, time zone and escalation. During an incident at 03:00 local time, who answers, in what language, and how fast does it reach an engineer who can change behaviour rather than open a ticket?
  • Interfaces you already run. Which flow telemetry versions are consumed, which BGP features are supported, is there an API, does it integrate with your existing authentication, and does it export to your existing logging and monitoring?

Scoring the tender

Two habits separate a tender that selects well from one that selects the best-written proposal.

The first is acceptance testing with your own traffic profile, not with the vendor’s demonstration. Mirror a representative slice of real production traffic, run the attack tests at the smallest frame size, enable the countermeasure set you intend to run in production rather than a minimal set, and measure both the mitigation behaviour and the false-positive effect on the legitimate traffic in the mirror. A platform’s numbers with everything switched on are the only numbers that describe what you have bought.

The second is weighting the criteria before the bids arrive. If the weights are decided afterwards, they will be decided by the proposals. A defensible weighting for an operator gives real mass to the tenancy and reporting layer, to the asymmetry answers, to the diversion and withdrawal mechanics, and to the supply and dependency questions — and treats headline throughput as a threshold to pass rather than a score to maximise, because throughput above your own edge capacity buys you nothing.

Ask, finally, for two references from operators of comparable size in comparable regulatory conditions, and ask them one question that proposals never answer: what did the platform do the first time it produced a serious false positive, and how long did it take to correct.

The specification checklist to lift into your RFP

Copy this into the tender and require an answer to each item in writing.

Capacity and performance

  1. Mitigation capacity in bits per second and packets per second, with the countermeasure set enabled that the operator intends to run.
  2. Behaviour and degradation mode when the packet-rate limit is reached.
  3. Backhaul requirement from edge routers to the mitigation platform.
  4. Number of tenants supportable under simultaneous mitigation, and the resource model behind that number.
  5. Performance figures validated by acceptance testing on the operator’s own mirrored traffic at the smallest realistic frame size.

Routing and diversion

  1. BGP session model for route injection, including authentication, maximum-prefix limits, prefix-length filters and community tagging of diversion routes.
  2. Behaviour of announced diversion routes on controller failure, restart or loss of management connectivity, including whether routes expire automatically.
  3. Supported return-path mechanisms, and whether encapsulation is performed in hardware.
  4. Confirmation that the intended diversion announcement remains valid under RPKI origin validation, with the ROA maxLength implications documented.
  5. RTBH signalling support, and the accepted community list, accepted prefix lengths and committed response times of each upstream provider.
  6. Support for RFC 8955 and RFC 8956, with the vendor-confirmed list of match and action combinations that run in hardware on the operator’s line-card inventory.
  7. FlowSpec rule cap, automatic expiry, per-tenant scoping, audit trail and kill switch.
  8. Support for DOTS signalling (RFC 8811, RFC 9132, RFC 8783) towards customers with their own on-premises devices.

Detection and inspection

  1. Flow telemetry versions consumed, and the export-timer assumptions that determine the floor on detection latency.
  2. Per-tenant baseline learning, with the ability to audit and exclude contaminated learning data.
  3. The explicit list of countermeasures that require symmetric traffic visibility.
  4. Default behaviour when a handshake is never observed.
  5. Load distribution model within the cluster (per-flow required), and fragment handling across cluster members.
  6. State model, state-table capacity, exhaustion behaviour and per-tenant state limits.

Multi-tenancy and reporting

  1. Separate baseline, threshold set, countermeasure order and allow-list per tenant on shared hardware.
  2. Policy templates, so that tenant count does not translate into hand-written configurations.
  3. Documented resource partitioning between tenants: state, rules and processing budget.
  4. Per-tenant portal restricted to that tenant’s prefixes, with attack timeline, vector breakdown, dropped-versus-passed volume, export, and role-based access control.
  5. API for tenant-side integration, and white-labelling of the portal and reports.

Commercial, supply and lifecycle

  1. Whether core detection functions without reachability to a vendor cloud.
  2. Exactly which capabilities licence expiry disables.
  3. Export-licensing and sanctions exposure affecting supply, support or payment.
  4. Spares location, RMA commitment and lead times for the operator’s region.
  5. Support language, hours, escalation path and time to reach an engineer with authority to change behaviour.
  6. The operator’s own published protection threshold, what happens above it, and how that threshold is reflected in the customer contract.

Sources and further reading

Every document below is a real standard or guidance text worth having on the table during the tender. No figures in this article are drawn from them; the single numerical example above is explicitly illustrative and must be recalculated from your own telemetry.

Flow telemetry. RFC 7011 for the IPFIX protocol and RFC 7012 for its information model; RFC 3954 documents NetFlow version 9 informationally; RFC 5475 for packet selection and sampling techniques and RFC 5476 for the PSAMP protocol. The current version of sFlow is an sFlow.org specification rather than an IETF standard — RFC 3176 documents only an earlier version, and informationally.

Routing and diversion. RFC 7454 (BCP 194) for BGP operations and security; RFC 3882 for destination-based remotely triggered blackholing; RFC 5635 for the source-based variant with uRPF; RFC 7999 for the BLACKHOLE community. RFC 2784 for GRE and RFC 2890 for its key and sequence-number extensions. RFC 4364 for MPLS-based L3VPN and VRF architecture.

FlowSpec. RFC 8955 for IPv4 and RFC 8956 for IPv6. The rule validation procedure and the defined action set are in those texts; confirm separately with your vendor which subset your own platform supports in hardware.

Routing security. RFC 6480 for the RPKI architecture, RFC 6482 for the ROA profile, RFC 6811 for prefix origin validation and RFC 9319 (BCP 185) for maxLength usage. RFC 2827 (BCP 38) for source address validation and RFC 3704 (BCP 84) for ingress filtering in multi-homed networks. NIST SP 800-189 provides an institutional framework covering interdomain routing security and DDoS mitigation. MANRS expresses much of the same material as operator commitments.

Standards-based mitigation signalling. RFC 8811 for the DOTS architecture, RFC 9132 for the signal channel and RFC 8783 for the data channel.

Analyst material. If your procurement process requires third-party validation, request the current network-security and DDoS-mitigation evaluations from the major analyst houses directly and read them against the criteria above; do not accept a vendor’s summary of a report you have not seen.

Frequently asked questions

What figure should an operator size DDoS mitigation capacity against?
Against your own edge, not against your customers' access circuits. The customer circuit limits how much of an attack reaches the customer; it says nothing about how much enters your network. The number that matters is what can physically arrive across the sum of your peering and transit ports. Two further limits are routinely forgotten: the backhaul capacity from your edge routers to the scrubbing cluster, and the packet-per-second budget of the mitigation platform. Your honest protection ceiling is the smallest of the three, not the largest.
Why does asymmetric routing matter when evaluating a mitigation platform?
Because an inspection engine that sees only one direction of a session cannot genuinely hold state. It cannot confirm that a TCP handshake completed, cannot follow sequence progression and cannot observe the session closing. In a multi-homed network, closest-exit routing makes asymmetry normal rather than exceptional, and out-of-path diversion makes it structural because only the ingress direction is diverted. Ask every bidder which specific countermeasures require symmetric visibility, and what the device does when it never sees the response.
Can I require FlowSpec from my upstream transit providers?
Usually not, or only for a narrow subset. Most transit providers do not accept FlowSpec rules from customers, and those that do restrict the match and action sets. The reasons are sound: a malformed rule propagates across the edge in seconds, hardware filter resources are finite, and the validation procedure in RFC 8955 constrains what can be expressed across AS boundaries. The one primitive that is available almost everywhere upstream is remotely triggered blackholing. Specify FlowSpec primarily for use inside your own AS, with a rule cap and automatic expiry.
What can an operator honestly commit to in a DDoS SLA?
Only what you measure and control: time from detection event to diversion announcement, time from diversion to the mitigation policy being applied, first-response time to a false-positive report, and the availability of the scrubbing platform and the customer portal as a service distinct from the customer's own availability. What cannot be committed is that every attack will be stopped, that there will be no false positives, that end-user experience will be unaffected, or that an attack larger than your edge capacity will be mitigated. Write the exclusions as plainly as the commitments.
Why is multi-tenancy the criterion that turns mitigation into a product?
Because without it you have network hygiene rather than something you can invoice. Selling the service requires three capabilities on shared hardware: a separate baseline and policy per customer, isolation so that an attack on one tenant cannot degrade another's protection by consuming shared resources, and per-tenant visibility so each customer sees only their own prefixes in a report they can export and show to their own auditors. The customer never sees your detection engine; they see the report.
Should the RFP ask for bits per second or packets per second?
Both, in writing, and treat the packet rate as the binding one. Attacks built from small packets reach the platform's packet-processing limit long before they approach its advertised bit rate, so a figure quoted only in gigabits describes the easiest case. Run acceptance testing at the smallest frame size the platform will realistically see, and require the vendor to state the packet rate at which the quoted mitigation behaviour still holds — with the same countermeasures enabled that you intend to run in production.
How should the tender handle vendor and supply-chain dependency?
Ask what still works if the vendor relationship stops. Hardware generally keeps forwarding and locally trained detection keeps working, but licence renewal, firmware and security patches, support, RMA and spare parts all depend on a live commercial relationship, and any cloud-delivered threat feed goes stale. The order matters more than the list: a product whose core detection depends on a vendor-operated cloud converts a slow degradation into an outage. Require the answer in writing, along with regional spares holding and lead times.
Can an operator with a small NOC realistically build its own scrubbing capacity?
It depends less on capacity than on how many consoles the build ends up needing. What makes it workable for a lean team is a single policy model carrying separate protection profiles and separate reporting per tenant on the same hardware, an API so technical customers pull their own data, and detection that keeps deciding when a vendor feed does not arrive. More than one appliance is built that way and they differ in emphasis: Corero SmartWall leans on automatic sub-second inline mitigation with minimal operator intervention, HARPP DDoS Mitigator on reaching its verdicts on the operator's own infrastructure with per-customer profiles on shared hardware. None of it moves the ceiling, which remains the smallest of your peering and transit edge, your backhaul and your packet-rate budget.

Published: August 2026

This guide is updated as vendors release new models and pricing. How we compare vendors