About
Envoy Latency and Fault Distribution Simulation
An Envoy dynamic module upstream HTTP filter written in Go that injects latency and fault responses based on configurable percentile distributions.
This is similar to Envoy's built-in fault injection filter but adds support for percentile-based latency distributions with per-status-code weighting — allowing you to simulate realistic endpoint behavior including error rates, latency profiles, and load-dependent degradation.
Key differentiator: By operating as an upstream filter on the cluster, this module measures the actual upstream response time and only injects the remaining delay needed to reach the target distribution value. If the upstream is already slower than the target, no additional delay is added.
Features
- Percentile-based latency injection: Configure latency distributions using flexible percentile notation (
p0.0,p50.0,p99.9,p100.0) - Per-status-code distributions: Define different latency profiles for different HTTP status codes (e.g., 200s are fast, 503s are slow)
- Resolution-based weighting: Use
resolutionas both the relative weight for status selection and the number of latency samples per status in stateful mode - Load-based behavior: Configure different response profiles based on the current in-flight request count with smooth grey-zone transitions
- Grey zone penalties: Model degradation with spike detection, penalty multipliers, and recovery rates
- Route matching: Apply different fault configurations via prefix/exact path matching and header matching
- First-match routing: Endpoints are evaluated in order; first match wins
- Upstream-aware timing: Measures actual upstream latency and only adds the remaining delay to reach the target — avoids over-delaying when the upstream is naturally slow
Comparison with Envoy's Built-in Fault Filter
| Feature | Built-in Fault Filter | This Module |
|---|---|---|
| Fixed delay | ✅ | ✅ (flat distribution) |
| Percentile distributions | ❌ | ✅ |
| Per-status-code distributions | ❌ | ✅ |
| Load-based degradation | ❌ | ✅ |
| HTTP abort | ✅ | ✅ |
| gRPC abort | ✅ | (planned) |
| Header-controlled faults | ✅ | ❌ (route-based instead) |
| Response rate limiting | ✅ | ❌ |
| Per-route configuration | Via per-route config | Built-in route matching |
| Runtime configuration | ✅ | ❌ |
| Precomputed per-status latency samples | ❌ | ✅ (stateful distribution) |
Configuration
Please find an overview of the possible fields below, followed by an actual example of how to configure the filter as a per cluster http filter.
Configuration Fields
Filter-level configuration retains endpoints[] for launchers that cannot express Envoy route
configuration. When this filter is configured through Envoy's typed_per_filter_config, the
per-route value must use the direct behavior shape: responses or load_based at the top level.
endpoints[] is rejected by the per-route configuration constructor because route matching is
performed by Envoy.
Note
Per-route configuration is not currently accessible through the boe CLI. It requires direct
control of the Envoy bootstrap configuration to add typed_per_filter_config to a route.
The sampled status is authoritative: a status different from upstream is returned as a local
response after any remaining configured delay. Matching statuses retain the upstream body and
headers. Configure local_response per response distribution to supply the forced response body
and headers. Omitted bodies are empty for sampled statuses below 400 and generated fault text for
statuses 400 and above; statuses 204, 205, and 304 must have an empty body. Envoy may remove
Content-Type for empty bodies and bodyless statuses.
| Field | Description |
|---|---|
endpoints |
Array of endpoint configurations. First match wins. |
endpoints[].match |
Optional route selector; omitted or empty matches all requests |
endpoints[].match.prefix |
Match requests whose path starts with this prefix |
endpoints[].match.exact |
Match requests with exactly this path |
endpoints[].match.headers |
Array of header match conditions (all must match) |
endpoints[].responses |
Array of status-code distributions (weighted by resolution) |
endpoints[].responses[].status |
Terminal HTTP status code (200-599) |
endpoints[].responses[].local_response |
Optional body and response header pairs for a forced response |
endpoints[].responses[].local_response.body |
UTF-8 body; omitted means empty below 400 and generated fault text at 400 and above |
endpoints[].responses[].local_response.headers |
Optional header pairs; repeated names are retained and Envoy calculates Content-Length |
endpoints[].responses[].resolution |
Weight for status selection AND number of pre-computed samples |
endpoints[].responses[].distribution |
Percentile-to-duration mapping |
endpoints[].load_based |
Load-sensitive behavior configuration |
endpoints[].load_based.healthy |
Behavior below the healthy in-flight request threshold |
endpoints[].load_based.tipping_point |
Behavior above the tipping point in-flight request threshold |
endpoints[].load_based.grey_zone |
Transition parameters between healthy and tipping |
diagnostic |
Include the diagnostic x-fault-worker-index response header; defaults to false |
probability_distribution |
Select latency sampling mode: random interpolation (stateless) or precomputed latency samples (stateful). Status selection remains weighted-random in both modes. |
Percentile Keys
Percentile keys use the format p<value> where value is between 0 and 100:
| Key | Quantile |
|---|---|
p0.0 |
0th percentile (minimum) |
p25.0 |
25th percentile |
p50.0 |
50th percentile (median) |
p75.0 |
75th percentile |
p90.0 |
90th percentile |
p95.0 |
95th percentile |
p99.0 |
99th percentile |
p99.9 |
99.9th percentile |
p99.99 |
99.99th percentile |
p100.0 |
100th percentile (maximum) |
Distribution values must be non-decreasing (a higher percentile cannot have a shorter duration).
Grey Zone Configuration
| Field | Description |
|---|---|
penalty_base |
Base latency penalty at full grey zone position (e.g., "50ms") |
spike_threshold |
Grey zone position (0-1) above which spike behavior activates |
spike_penalty_duration |
Recovery window after load drops below the spike threshold (e.g., "2s") |
spike_penalty_multiplier |
Multiplier applied to base penalty during spikes |
recovery_rate |
Rate at which spike penalty decays (0-1) |
How It Works
Status Code Selection
Each endpoint has one or more response entries with a resolution that serves as both:
- Weight: The probability of selecting that status code (proportional to total resolution)
- Accuracy: The number of pre-computed latency samples for that status code's distribution
For example, with resolution: 900 for status 200 and resolution: 100 for status 503, each
request has a 90% chance of selecting status 200 and a 10% chance of selecting status 503. The
selected status is returned even when it differs from upstream; observed status counts approach
the configured weights over time.
Latency Distribution
The stateful probability distribution is inspired by distribution-calculator. Given a set of percentiles, it:
- Precomputes latency samples based on
resolutionby interpolating between percentile boundaries - Shuffles and serves them in random order
- The samples approximate the configured percentiles; status selection remains independently weighted-random, and observed latency cannot be lower than upstream latency
Load-Based Behavior
When load_based is configured, the process-global active-request count is passed as the current load value:
- At or below
healthy.threshold_in_flight: Uses the healthy response distribution - At or above
tipping_point.threshold_in_flight: Uses the tipping point distribution - Between the two (grey zone): Probabilistically mixes between healthy and tipping based on position, with optional penalty
Grey Zone Transitions
In the grey zone, the filter:
- Calculates position as
(currentInFlight - healthyThreshold) / (tippingThreshold - healthyThreshold)(0.0 to 1.0) - Selects healthy or tipping distribution proportionally to position
- Adds a base latency penalty scaled by position
- While position is at or above
spike_threshold, keeps the spike multiplier active - Starts recovery on the first observed drop below the threshold. For
spike_penalty_duration, the penalty isbasePenalty * spike_penalty_multiplier * (1 - elapsed / spike_penalty_duration * recovery_rate), then returns to the base penalty. A renewed spike resets recovery
Spike state follows load observations in every tier. The tipping tier sustains a spike; the healthy tier starts or advances recovery. Additional penalties apply only in the grey zone.
Response Headers
The filter adds response headers to indicate what was injected:
| Header | Description |
|---|---|
x-fault-injected-delay |
Target duration from the distribution (e.g., "52.3ms") |
x-fault-actual-upstream |
Actual time the upstream took to respond |
x-fault-added-delay |
Additional delay injected (target - upstream); omitted when no delay was added |
x-fault-requests-in-flight |
Process-global in-flight matched request count observed at request entry, before the request is added |
x-fault-worker-index |
Envoy worker index that made the fault decision; only included when diagnostic is true |
x-fault-injected |
"response" when a sampled status below 400 overrides upstream, otherwise "abort" |
x-fault-status |
The status code selected by the distribution |
x-fault-upstream-status |
The original upstream status code, including when replaced by the sampled status |
For matched responses, these attributes are also recorded on the active span when one is available,
using the fault. prefix (for example, fault.status and fault.upstream-status). This preserves
visibility into backend failures even when the sampled response is successful.
Performance considerations
The stateful and stateless probability distributions have different CPU,
memory, and concurrency characteristics. The benchmark sources in
performance_bench_test.go, endpoint_bench_test.go,
internal/fault/performance_bench_test.go, and performance/ reproduce the
measurements below.
Use stateful when CPU efficiency and precomputed latency samples are
important; it remains the default. Use stateless for high-concurrency,
throughput-sensitive configurations where lower contention, tail latency, and
resolution-independent memory matter more, accepting higher CPU use. Load-based
configurations show little difference because their outer mutex dominates both
modes.
Usage Examples
Delay all success responses and introduce a 1% failure rate that returns quickly
Wrap a succesfully and fast responding mock or stub service and make it a one matching the production performance and error ratio.
CONFIG=$(cat <<-END
{
"endpoints": [
{
"match": {
"prefix": "/"
},
"responses": [
{
"status": 200,
"resolution": 1000,
"distribution": {
"p0.0": "30ms",
"p100.0": "500ms"
}
},
{
"status": 503,
"resolution": 10,
"distribution": {
"p0.0": "3ms",
"p100.0": "5ms"
}
}
]
}
]
}
END
)
boe run \
--extension dynamic-fault-injection \
--log-level dynamic_modules:debug \
--config "${CONFIG}"
❯ curl -v http://localhost:10000/status/200
> GET /status/200 HTTP/1.1
> Host: localhost:10000
> User-Agent: curl/8.7.1
> Accept: */*
>
< HTTP/1.1 200 OK
< date: Fri, 02 Oct 2026 19:11:25 GMT
< content-type: text/html; charset=utf-8
< content-length: 0
< server: envoy
< access-control-allow-origin: *
< access-control-allow-credentials: true
< x-envoy-upstream-service-time: 113
< x-fault-injected-delay: 326.1ms
< x-fault-actual-upstream: 114.130167ms
< x-fault-added-delay: 211.969833ms
< x-fault-status: 200
< x-fault-upstream-status: 200
< x-fault-requests-in-flight: 0
# About 1% of requests sample the 503 response instead and get a local reply
❯ curl -v http://localhost:10000/status/200
> GET /status/200 HTTP/1.1
> Host: localhost:10000
> User-Agent: curl/8.7.1
> Accept: */*
>
< HTTP/1.1 503 Service Unavailable
< content-type: text/plain
< x-fault-injected: abort
< x-fault-injected-delay: 4ms
< x-fault-actual-upstream: 182.419083ms
< x-fault-status: 503
< x-fault-upstream-status: 200
< x-fault-requests-in-flight: 0
< content-length: 24
< date: Fri, 02 Oct 2026 19:12:39 GMT
< server: envoy