Resource library

QA How-To

How to Fix k6 "thresholds have been crossed" (Exit Code 99)

Fix k6 thresholds have been crossed: diagnose exit code 99, find the failing metric, repair latency or errors, and verify your load-test gate safely in CI.

19 min read | 3,162 words

TL;DR

Exit code 99 means one or more configured k6 thresholds evaluated false. Find the failed row in the summary, repair the corresponding latency, HTTP failure, check, load-profile, or configuration issue, and rerun at the original load without suppressing the exit status.

Key Takeaways

  • Exit code 99 means at least one k6 threshold failed; inspect the failed metric before changing limits.
  • Separate latency, HTTP failure, and check thresholds because each reports a different problem.
  • Compare measured p95, request count, and achieved load with the exact SLO and traffic profile.
  • Use stable request tags when endpoints need different latency or availability budgets.
  • Check dropped iterations and container networking before blaming the application for a CI-only failure.
  • Keep the nonzero exit code in CI and prove the gate fails with a controlled strict-threshold run.

To fix k6 thresholds have been crossed errors, start with the load test or cloud run that reported a breached pass/fail limit. Exit code 99 is k6's threshold failure signal. Read the failed metric and its measured value before changing a limit.

Thresholds have been crossed

That exact wording is used by a cloud threshold failure. A local run may instead print a threshold breach message naming the failed metrics, followed by the same exit code. The message alone cannot tell you whether latency, HTTP failures, failed checks, or a custom metric caused the result.

TL;DR

Find the red cross in the THRESHOLDS part of the k6 summary. Compare the observed value with the expression beside it, then inspect the corresponding requests. If http_req_duration fails, investigate server time and the load profile. If http_req_failed fails, inspect response status codes and connectivity. If checks fails, inspect the assertion and response body. Preserve exit code 99 in CI so a real regression still blocks the job.

k6 run --summary-mode=full thresholds.js
echo $?

The command assumes you have installed k6 and created thresholds.js as shown below. A result of 0 means all configured thresholds passed; 99 means at least one failed. Grafana's threshold documentation defines the pass/fail behavior, and its threshold validation guide documents the exit codes. Keep the first failing run's summary. A later green run at a different traffic level does not explain the original failure.

What the Error Actually Means

A threshold is an expression over an aggregated metric, such as p(95)<500 for request duration in milliseconds or rate<0.01 for the proportion of failed HTTP requests. k6 evaluates each configured expression against the samples produced by the run. One false expression is enough to mark the test failed. Checks are different: a failed check() records an unsuccessful assertion, but it affects the process result only if you also define a threshold on checks or another affected metric.

The error is therefore a useful quality gate, not proof that k6 itself crashed. http_req_duration is a Trend; p(95) is its 95th percentile. http_req_failed and checks are Rates. For http_req_failed, a lower rate is better. For checks, a higher rate is better. A displayed rate of 1.00% does not satisfy rate<0.01, because the operator is strict. The numerical limit in each example here is illustrative; use your service's actual SLO and representative load when choosing production criteria.

A request can pass a status check and still be slow. Conversely, a fast error response can keep the duration percentile low while raising http_req_failed. A global latency threshold also mixes every HTTP request unless you filter it with tags. Review the metric name, selector, expression, actual value, request count, and run profile together. The k6 thresholds and checks guide explains the two mechanisms, while reading p95 and p99 latency helps interpret a percentile breach.

To reproduce the commands, start a disposable local server in one terminal, then put this complete file in the directory where you run k6. The local server is a diagnostic target only, not a substitute for your staging service.

python3 -m http.server 8000 --bind 127.0.0.1
// thresholds.js
import http from 'k6/http';
import { check, sleep } from 'k6';

export const options = {
  vus: 2,
  duration: '20s',
  thresholds: {
    http_req_duration: ['p(95)<500'],
    http_req_failed: ['rate<0.01'],
    checks: ['rate>0.99'],
  },
};

export default function () {
  const base = __ENV.BASE_URL || 'http://127.0.0.1:8000';
  const response = http.get(`${base}/`, {
    tags: { endpoint: 'home' },
    timeout: '5s',
  });
  check(response, { 'home returns 200': (r) => r.status === 200 });
  sleep(1);
}

Run k6 run thresholds.js while the server is listening. The sample output depends on your machine, but you should see three threshold rows and echo $? should print 0. If the server is unavailable, the failure is expected: the HTTP error rate and check rate expose the broken target. The k6 load testing tutorial covers the basic runner setup.

Root-Cause Decision Table

Symptom Likely root cause Fix to verify
http_req_duration p95 crosses its limit while HTTP failures stay low Slow application path or a load level beyond capacity Compare endpoint and server timings at the intended load
http_req_failed rate rises Error status, connection failure, DNS/TLS issue, or incorrect expected-status policy Inspect statuses, request errors, and target URL
checks rate fails but HTTP failures stay low Functional assertion fails on a successful HTTP status Inspect the response content and assertion
Global p95 fails but the critical endpoint passes Slow setup or secondary request is included Add an endpoint tag and threshold the intended subset
One failure in a tiny run breaches a strict rate Too few samples for a stable decision, or an unrealistic limit Increase representative sample volume and calibrate to the SLO
dropped_iterations grows or CI alone fails Generator cannot maintain requested arrival rate, or CI reaches a different target Correct executor capacity and target configuration
A run aborts during startup abortOnFail evaluates during warmup Delay abort evaluation only when warmup is outside the SLO window

1. Fix k6 Thresholds Have Been Crossed When Latency Regresses

First decide whether the slow requests are genuine. Run the same profile against the same environment and compare its p95 with a last known good run. The http_req_duration metric measures HTTP request duration, not the full sleep(1) in the sample iteration. A failed p(95)<500 means the measured p95 was at least 500 milliseconds. Raising the limit to 1000 milliseconds would make the gate green, but it would hide a regression if the SLO still says 500.

Run a full summary and record the request count, rate, and p95. Then export individual data points for diagnosis. Use a bounded diagnostic run, because JSON output grows with traffic.

BASE_URL=http://127.0.0.1:8000 k6 run --summary-mode=full --out json=/tmp/k6-points.json thresholds.js
jq -r 'select(.type == "Point" and .metric == "http_req_duration") | [.data.time, .data.value, .data.tags.endpoint] | @tsv' /tmp/k6-points.json | head

The jq rows show sampled duration, timestamp, and the endpoint tag assigned in thresholds.js. For a real service, point BASE_URL at a staging host you own and correlate slow timestamps with server traces, database query time, CPU saturation, and upstream calls. Change one bottleneck, rerun at the original load, and verify the same threshold passes without lowering traffic. The performance bottleneck guide goes deeper on tracing a hot path. Do not infer a root cause from p95 alone: a slow database, queueing at a connection pool, and a slow remote dependency can have the same summary value.

2. Fix k6 Thresholds Have Been Crossed When HTTP Errors Spike

http_req_failed is based on whether k6 considers each response expected. By default, status codes from 200 through 399 count as expected. A timeout, connection error, or unexpected status can increase the failure rate. A check() on status === 200 is narrower than that default: a redirect could count as an expected HTTP response yet fail the check. That difference explains why the http_req_failed and checks rows sometimes disagree.

Start with one request so debugging output stays readable. Stop the local server to reproduce a connection failure, then restart it to verify recovery. For a staging issue, substitute its URL and inspect the response status, redirects, DNS, and authentication path.

BASE_URL=http://127.0.0.1:8000 k6 run --vus 1 --iterations 1 --http-debug thresholds.js

This command can print headers or response details. Use it only with a safe diagnostic target, because request data can contain sensitive headers. If a non-2xx status is intentionally successful for one request, set a per-request response callback rather than globally treating all errors as expected. For example, a status endpoint designed to return 204 can use responseCallback: http.expectedStatuses(204) in its request parameters. Grafana documents expectedStatuses and the per-request callback parameter. Verify by rerunning the one-request command and checking that the HTTP failure row falls while the response still matches the endpoint contract. Never mark a real 500 as expected to make a dashboard green.

3. Repair a Failed Check Without Masking HTTP Failures

A checks threshold fails when too many check() predicates return false, even if every request completed and k6 marked its status expected. This often happens when a page or API returns 200 with a login screen, an error envelope, stale content, or a missing field. Inspect the actual response before broadening the predicate. A check of only status === 200 cannot detect a broken body; a content check can.

The following separate file targets the local Python server and checks both status and a string that server returns in its directory listing. Save it as checks.js. Keep the HTTP failure threshold so a successful body assertion does not replace transport monitoring.

// checks.js
import http from 'k6/http';
import { check } from 'k6';

export const options = {
  vus: 1,
  iterations: 3,
  thresholds: {
    http_req_failed: ['rate<0.01'],
    checks: ['rate>0.99'],
  },
};

export default function () {
  const response = http.get('http://127.0.0.1:8000/');
  check(response, {
    'status is 200': (r) => r.status === 200,
    'listing has heading': (r) => r.body.includes('Directory listing for'),
  });
}
k6 run checks.js
echo $?

With python3 -m http.server still running, both checks should pass and the exit code should be zero. If you serve a different page, replace the body predicate with a stable feature of that page. To diagnose a failing real API check, log the relevant status and a safe excerpt of the body during a one-VU run, then remove that diagnostic logging. The API performance testing tutorial shows how to combine response validation with load.

4. Scope an Aggregate Threshold to the Endpoint You Care About

The sample http_req_duration threshold combines all HTTP requests in the run. An authentication call, static asset, or health probe may have a very different latency budget from the checkout API. If they share one global p95, the summary cannot show whether checkout met its own SLO. Give each request a stable tag and place thresholds on the tagged sub-metrics. Avoid high-cardinality values such as user IDs or raw query strings in tags.

Save this complete example as tagged.js. It uses two calls to the same local server so you can verify the selector mechanics before substituting real endpoints. Replace the URLs and limits with service-specific values only after collecting a baseline.

// tagged.js
import http from 'k6/http';

export const options = {
  vus: 2,
  duration: '20s',
  thresholds: {
    'http_req_duration{endpoint:home}': ['p(95)<500'],
    'http_req_duration{endpoint:listing}': ['p(95)<500'],
    http_req_failed: ['rate<0.01'],
  },
};

export default function () {
  const base = 'http://127.0.0.1:8000';
  http.get(`${base}/`, { tags: { endpoint: 'home' } });
  http.get(`${base}/?view=listing`, { tags: { endpoint: 'listing' } });
}
k6 run --summary-mode=full tagged.js

Verify that both tagged duration thresholds appear in the summary and inspect each p95 separately. In a real workflow, use http_req_failed{endpoint:checkout} if checkout errors need their own limit as well. A missing tag value can leave you with a threshold that has no useful samples, so check the request count and exported point tags when a row looks suspicious. The k6 scenarios and executors guide is useful when endpoints are also separated into different traffic profiles.

5. Calibrate Strict Rates and Small-Sample Percentiles

A rate threshold has exact mathematics. With 100 requests, one failed request yields an HTTP failure rate of 0.01. That fails rate<0.01, because 0.01 is not less than 0.01. If the SLO permits at most 1%, rate<=0.01 expresses that boundary. If the SLO requires fewer than 1%, leave the strict operator and fix the failed request. Do not change the operator because a single run was inconvenient.

Short smoke runs also give noisy percentiles. With only a few requests, one slow connection setup or cold cache can move p95 sharply. First match the test's workload, data, and warmup assumptions to the service-level objective. Then collect enough requests for a meaningful distribution. A larger sample does not excuse a failing availability budget, but it makes the estimate less sensitive to one observation. Record both the sample count and percentile in review.

For a controlled boundary check, make a copy of thresholds.js named strict.js and change only its http_req_duration expression to p(95)<1. Run the copied file against the local server, capture echo $?, then restore the intended threshold in your real test. A failed run should return 99; the original thresholds.js should return zero while the local server is healthy. This verifies that CI can detect a breach without weakening the production gate. Grafana's threshold validation exercise uses the same pass/fail pattern. For deeper interpretation of sample distributions, read percentile latency p95 and p99.

6. Correct an Unsustainable Arrival Rate or Weak Load Generator

A test can miss its target traffic before the service itself is saturated. With an arrival-rate executor, k6 schedules iterations at a target rate. If there are too few available virtual users or the CI runner is CPU-starved, it can drop iterations. A passing latency threshold under reduced throughput would then be misleading. Add an explicit dropped_iterations threshold and review achieved request rate alongside response time.

Save this diagnostic profile as arrival.js. The rates and VU pool are illustrative and deliberately small for the local server. Increase them only in an environment designed to take load.

// arrival.js
import http from 'k6/http';

export const options = {
  scenarios: {
    steady: {
      executor: 'constant-arrival-rate',
      rate: 2,
      timeUnit: '1s',
      duration: '20s',
      preAllocatedVUs: 5,
      maxVUs: 10,
    },
  },
  thresholds: {
    dropped_iterations: ['count<1'],
    http_req_failed: ['rate<0.01'],
    http_req_duration: ['p(95)<500'],
  },
};

export default function () {
  http.get('http://127.0.0.1:8000/');
}
k6 run --summary-mode=full arrival.js

Verify dropped_iterations stays at zero and that roughly the requested iteration rate was achieved. If it rises, check load-generator CPU, network capacity, test-side JavaScript, and VU allocation before claiming the application passed. Increase preAllocatedVUs only when the generator can support it, then rerun the same workload. The load testing guide explains why traffic shape matters; the k6 scenarios guide covers executor selection.

7. Handle Warmup and Early Abort Deliberately

abortOnFail can end a run before enough samples arrive. A cache cold start or a single early error may breach the threshold temporarily even though the steady-state window would pass. Conversely, if startup behavior is part of the user-facing SLO, excluding it would conceal a real problem. Decide which period your test measures, then encode that decision explicitly.

Use the long threshold form when a critical failure should stop a costly run. This complete warmup.js example delays abort evaluation for ten seconds; it does not waive the final threshold. The same local server can verify the syntax and final result.

// warmup.js
import http from 'k6/http';
import { sleep } from 'k6';

export const options = {
  vus: 2,
  duration: '30s',
  thresholds: {
    http_req_duration: [
      { threshold: 'p(95)<500', abortOnFail: true, delayAbortEval: '10s' },
    ],
    http_req_failed: ['rate<0.01'],
  },
};

export default function () {
  http.get('http://127.0.0.1:8000/');
  sleep(1);
}
k6 run warmup.js
echo $?

Expect zero on a healthy local target. If you stop the server, the HTTP failure threshold still fails at the end, regardless of the delayed latency abort. Grafana notes that cloud threshold evaluations occur periodically, so an abortOnFail cloud run may stop later than a local run. Do not use the delay as a blanket workaround for recurring warmup problems; instrument cache fill, connection creation, or deployments if those are what users experience.

8. Fix CI or Docker Environment Drift Without Swallowing Exit 99

A job may reach a different host than your laptop, use missing credentials, or run with far less CPU. Keep target selection explicit. The baseline script reads BASE_URL, so CI should set that value to a dedicated test environment rather than relying on the local default. From the runner, probe the target before load and make sure the expected response is visible. Then run k6 and preserve its process status.

curl -i --max-time 5 "$BASE_URL/"
k6 run --summary-mode=full thresholds.js

For Docker, use the installed k6 image tag that matches your approved runner version. Do not assume that 127.0.0.1 inside a container refers to a server on the host. Put both services on a Docker network or use a routable test URL, then supply BASE_URL explicitly. This example assumes BASE_URL is already a reachable staging URL and K6_IMAGE contains a pinned image reference chosen by your team.

docker run --rm -e BASE_URL="$BASE_URL" -v "$PWD":/scripts -w /scripts "$K6_IMAGE" run --summary-mode=full thresholds.js

Verify curl reaches the right environment, then compare Docker and host summaries at the same profile. If Docker alone has failures, inspect DNS, TLS trust, container resource limits, and authentication before changing thresholds. In a shell pipeline, k6 run thresholds.js | tee k6.log can hide the k6 status unless the shell enables pipefail. In Bash, use set -o pipefail before the pipeline, or run k6 without tee. Exit 99 is a failed quality gate and should remain nonzero. The CI/CD troubleshooting guide provides broader runner checks.

How to Verify the Fix

Run the original test, target, and load profile after correcting the identified cause. Confirm all threshold rows have a pass mark and that echo $? prints 0. Compare the failed metric's new value, request count, achieved throughput, and relevant server signals with the failed run. For a performance regression, the fix is credible only if the original SLO passes at the original load. For a configuration error, show that the target, tag, or expected-status rule now reflects the actual contract.

BASE_URL=http://127.0.0.1:8000 k6 run --summary-mode=full thresholds.js
status=$?
printf 'k6 exit status: %s\n' "$status"
exit "$status"

Use the local URL for the demo and replace it with the authorized staging target in your pipeline. To prove that the gate still fails, run a deliberately strict copy such as strict.js from section 5 and check for exit 99. Do this in a controlled environment, then keep the realistic threshold in the maintained test. A green run after removing the threshold does not verify a fix.

Prevent It From Coming Back

Keep threshold expressions in code review with the SLO they enforce, the traffic profile they assume, and the environment where they apply. Track a known-good baseline with request count and p95, not just a screenshot of a green check. Separate critical endpoints with stable tags so unrelated requests do not rewrite a quality gate's meaning. Monitor both latency and failures; a service can improve p95 by rejecting requests quickly.

In CI, retain the complete k6 summary and the job's exit status. Run a small diagnostic profile on pull requests and a representative load profile in a controlled performance environment. Compare achieved arrival rate with the requested rate. When you change data volume, authentication, routing, or server capacity, remeasure before editing the limit. The k6 performance engineering guide connects these checks to longer-running capacity work.

Interview Questions and Answers

Q: Why does k6 exit with 99? A configured threshold evaluated false. Name the exact metric and expression from the summary before proposing a fix.

Q: Does a failed check() automatically fail the run? No. Checks record pass/fail samples; a threshold such as checks: ['rate>0.99'] makes their aggregate a run gate.

Q: What does p(95)<500 measure? It requires the 95th percentile of a Trend metric such as http_req_duration to be below 500 milliseconds. It does not require every request to be below 500.

Q: Why can http_req_failed disagree with a status check? k6's default expected-status range is 200 through 399, while a user check may require exactly 200 or a specific body field.

Q: When should you use abortOnFail? Use it when continuing a severely failing test wastes time or risks stressing an unhealthy system. Add a deliberate delay if startup samples should not trigger early termination.

Q: How do you debug a threshold that passes locally but fails in CI? Match the URL, load profile, data, and runner capacity, then compare failed metrics and achieved throughput. Check Docker networking and whether the CI runner has credentials and enough resources.

Common Mistakes

  • Raising a latency limit solely to clear exit 99 without checking the SLO or the server-side regression.
  • Removing thresholds or adding || true to CI, which hides the very condition the test was built to detect.
  • Treating checks and http_req_failed as identical even when their success rules differ.
  • Measuring one mixed global percentile for endpoints with distinct latency budgets.
  • Reading p95 without request count, dropped iterations, and achieved traffic rate.
  • Running --http-debug at full load and flooding logs with sensitive response data.
  • Pointing a container at 127.0.0.1 while the service is outside that container.

Conclusion

To fix k6 "thresholds have been crossed," identify the failed threshold, inspect its metric under the intended load, correct the system or test definition that caused it, and rerun the original gate. Keep exit code 99 visible in CI. The right final evidence is a passing threshold at representative traffic, with enough summary detail to show why it passed.

Interview Questions and Answers

How would you investigate k6 exit code 99 in a release pipeline?

I would keep the failed job's summary and locate the threshold expression that evaluated false. Then I would compare its value, sample count, achieved throughput, and target environment with the last passing run. I would fix the service regression or the incorrect test definition, rerun the same workload, and retain the nonzero gate for future failures.

What is the difference between a k6 check and a threshold?

A check evaluates a predicate for an individual response or event and records pass/fail samples. A threshold evaluates an aggregate metric against a limit and determines the run result. To gate on checks, I would add an explicit expression such as `checks: ['rate>0.99']`.

Why can HTTP error rate pass while a 200-only check fails?

k6 considers 200 through 399 expected HTTP statuses by default, while a check can require exactly 200. A redirect can therefore count as an expected HTTP response yet fail the check. I would inspect the redirect chain and decide which status is valid for that endpoint.

What would you inspect after a p95 latency threshold fails?

I would confirm the load profile and number of requests, then separate the slow endpoint with a stable tag. I would correlate its slow timestamps with server traces, query time, saturation, and upstream latency. I would rerun the original SLO at the original load after addressing the identified bottleneck.

How do dropped iterations affect confidence in a k6 result?

Dropped iterations mean an arrival-rate scenario did not start all scheduled work. A latency pass under lower-than-requested traffic is not evidence that the service meets the intended load target. I would measure the generator's resources, allocate enough VUs, and rerun with a threshold on `dropped_iterations`.

When is a tagged threshold better than a global threshold?

I would use a tagged threshold when endpoints have different SLOs or a secondary request distorts a global aggregate. Each request gets a stable tag such as `endpoint:checkout`, and the threshold selects that tag. I would verify the sub-metric receives samples before trusting the result.

How can you test that CI will fail when a threshold is breached?

I would run a controlled copy of the script with an intentionally strict threshold in a safe environment. The expected outcome is exit 99 and a failed CI step. Then I would restore the approved SLO and verify the normal run exits zero.

Frequently Asked Questions

What does k6 exit code 99 mean?

It means a configured threshold failed. Inspect the THRESHOLDS summary or cloud run details for the metric, expression, and observed value. Preserve this status in CI because it is the intended failure signal.

Why does k6 say thresholds have been crossed when requests returned 200?

A successful HTTP status does not guarantee that latency or a body check met its limit. Look for a failed `http_req_duration` or `checks` row. A global duration threshold can also include other requests in the run.

How do I see which k6 threshold failed?

Run the script with `--summary-mode=full` and inspect the THRESHOLDS section for the failed expression and actual aggregate. In Grafana Cloud, inspect the run's threshold results. Keep the complete summary with the CI artifact.

Do failed k6 checks cause exit code 99?

A failed `check()` records a failed check sample. It causes exit 99 when a threshold on `checks` or another affected metric evaluates false. Without a threshold, failed checks alone do not define the run's pass/fail status.

Should I increase the p95 threshold to fix exit code 99?

Only if the old value was not the approved latency objective or the measured workload does not match the intended test. Otherwise investigate the slow path and rerun at the original load. Keep a record of the SLO decision when changing the limit.

Why does a k6 threshold fail in Docker but pass on the host?

The container may reach a different address, lack credentials or TLS trust, or have different CPU and network capacity. In particular, `127.0.0.1` inside Docker points to the container. Supply a routable test URL and compare request counts and achieved load.

What does abortOnFail change?

It lets k6 stop a test when the specified threshold fails before the planned end. `delayAbortEval` postpones early evaluation to gather initial samples. The threshold still needs to pass; these options do not make a failing result successful.

Related Guides