Resource library

QA How-To

k6 --summary-export JSON: Save and Compare Load Test Results

Use k6 summary export JSON to save aggregated load test results, inspect metrics, compare baselines, and avoid misleading performance regression claims.

22 min read | 2,815 words

TL;DR

Run `k6 run --summary-export baseline.json script.js` to save aggregated metrics, checks, and thresholds. Export a candidate with the same workload, inspect the JSON schema, then compare p95 latency, error rate, and delivered requests while preserving each run's k6 exit status.

Key Takeaways

  • Run k6 with --summary-export to save one aggregate JSON document after the test.
  • Inspect the export shape before parsing because machine-readable and legacy summaries differ.
  • Compare the same script, workload, target, and metric across baseline and candidate runs.
  • Use thresholds for each run and a separate comparison rule for changes across runs.
  • Check request count and failure rate before interpreting a latency improvement.
  • Prefer handleSummary() when designing a new long-lived reporting contract.

The k6 summary export JSON option saves one end-of-test document of aggregated load test results. Run k6 run --summary-export baseline.json script.js, preserve the file, and compare it with another run made under the same conditions. This guide builds a local target, exports two summaries, and reads their latency, failure rate, and request count with a small Python program.

A summary file is evidence only when you know how the test ran. A p95 latency number from one VU cannot be compared fairly with a p95 from a different scenario, target, or route mix. We will keep those inputs constant and show which details to record for a real release decision.

Grafana currently documents --summary-export as available but discourages new integrations based on it. Its k6 options reference recommends handleSummary() for flexible reporting. You will learn the flag because existing CI jobs use it, then see when to switch to a controlled custom summary.

What You Will Build

  • A loopback HTTP endpoint that receives five safe tutorial requests.
  • A k6 script with an HTTP check and thresholds that determine exit status.
  • baseline.json and candidate.json, each containing one end-of-test summary.
  • A Python comparator that rejects missing values and mixed summary schemas.
  • A repeatable handoff to a CI job that keeps the k6 exit code and JSON artifact.

The local run demonstrates commands and JSON shape. Five samples cannot prove application performance. In a release test, you would use an approved environment, enough traffic for useful statistics, and repeated runs to estimate normal variation.

Prerequisites

Install k6 v2.0.0 or newer with --summary-export support and Python 3.9 or newer. The documented k6 v2 command line supports the flag; use the current release available for your platform and record its exact output. Python's standard library is sufficient, so this exercise has no package pins. Use the official k6 installation guide if the executable is absent. Match any Docker image tag to the release you actually run rather than copying an unverified tag.

k6 version
python3 --version
k6 run --help | grep summary-export

Verify: all three commands succeed and the final line names summary-export. Keep a note of the actual versions alongside each future baseline. Open two terminal sessions in one empty working directory. The commands assume a Unix-like shell; on Windows, run the same k6 and Python commands in a compatible shell and adapt process management in the CI example.

For a meaningful comparison later, also prepare a stable target you are allowed to test, a written workload, and access to the previous run's artifact. Keep credentials outside exported files. Do not assume a JSON file alone describes the server build, script revision, or load generator.

Step 1: Start a Local HTTP Target

In terminal A, start Python's static file server. It returns a directory listing at /, enough for a repeatable HTTP 200 check. Binding to loopback keeps the demonstration on your computer.

python3 -m http.server 8000 --bind 127.0.0.1

Leave the server running. In terminal B, use Python to prove the endpoint is reachable before involving k6:

python3 -c "from urllib.request import urlopen; print(urlopen('http://127.0.0.1:8000/', timeout=5).status)"

Verify: the command prints 200, and terminal A logs a GET request. If port 8000 is occupied, choose another port in both the server command and the k6 script below. A connection refusal at this stage is a target setup problem, not a k6 export problem.

This server is intentionally simple. It cannot model a real API's database, caching, or concurrency behavior. Its value is isolation: when a file is missing or a field is unexpected, you can debug the export path without guessing whether an external service throttled the test. Do not use localhost results to claim release capacity.

Step 2: Write a k6 Test With Explicit Success Criteria

Save the following file as script.js. One virtual user runs five iterations, and each iteration makes one request. The check() records HTTP correctness. The thresholds cause a nonzero process exit when the check rate, failure rate, or latency limit is violated.

import http from 'k6/http';
import { check } from 'k6';

export const options = {
  vus: 1,
  iterations: 5,
  summaryTrendStats: ['avg', 'min', 'med', 'max', 'p(90)', 'p(95)'],
  thresholds: {
    checks: ['rate==1'],
    http_req_failed: ['rate==0'],
    http_req_duration: ['p(95)<5000'],
  },
};

export default function () {
  const baseUrl = __ENV.BASE_URL || 'http://127.0.0.1:8000';
  const response = http.get(`${baseUrl}/`, {
    tags: { operation: 'home' },
  });
  check(response, { 'home returns 200': (r) => r.status === 200 });
}

The five-second threshold is deliberately loose for this mechanics exercise. A real service threshold should come from an agreed requirement and a representative load profile. summaryTrendStats requests the percentile that the comparator will read. If you omit p(95) from that list, a summary consumer may not find the expected value. The operation tag will help if you extend the script to multiple routes; a global latency statistic can conceal one slow business-critical operation.

k6 run script.js

Verify: terminal A logs five GETs, and k6 reports five iterations, passing checks, and passing thresholds. Fix any failed check before exporting a baseline. Failed checks are measurements; the checks threshold is what makes them a CI gate. The k6 thresholds and checks guide shows how to define criteria for more complex workloads.

Step 3: Save the k6 Summary Export JSON Baseline

Run the same script with the output path supplied to --summary-export. The filename is relative to your working directory. k6 can replace an existing file, so use an immutable run identifier in shared automation; baseline.json is only a tutorial name.

k6 run --summary-export baseline.json script.js
python3 -m json.tool baseline.json > /dev/null
ls -lh baseline.json

Verify: the k6 run exits successfully, json.tool accepts the file, and ls reports a nonempty file. If a threshold fails, inspect the result even if the JSON exists. A written artifact means that the test reached summary generation, not that the service met its objective. Preserve the process exit code separately in a pipeline.

--summary-export records aggregate statistics at the end. It is different from k6 run --out json=points.json script.js, which streams line-delimited metric declarations and samples while the test runs. The official JSON output documentation describes that granular format. A line-delimited stream is useful for time windows and tags, but you cannot parse the entire file as a single JSON object.

Output Example command Data shape Good question
Summary export --summary-export baseline.json One JSON document of aggregates Did p95 or error rate change between runs?
Real-time JSON --out json=points.json One JSON object per line When did errors start, and which samples had a tag?
Custom summary handleSummary(data) Files or streams your function returns What stable report should another system consume?

A summary is compact and easy to archive. Its cost is lost timing detail: you cannot reconstruct a short spike from an aggregate percentile alone. A point stream gives that detail but consumes more storage and analysis work. Choose the export based on the question, and collect server metrics when diagnosing cause.

Step 4: Inspect the JSON Shape and Metric Units

k6 has legacy and newer machine-readable summary shapes. Newer output can expose a results.metrics array of named metric objects; older output commonly has a top-level metrics object. Verify the actual file before copying a JSON path from an example. The official k6 summary schema is the reference for the newer format.

python3 - <<'PY'
import json
with open('baseline.json', encoding='utf-8') as file:
    data = json.load(file)
print('Top-level keys:', sorted(data))
if isinstance(data.get('results'), dict):
    metrics = data['results'].get('metrics', [])
    print('Metric names:', [m.get('name') for m in metrics])
elif isinstance(data.get('metrics'), dict):
    print('Metric names:', list(data['metrics']))
else:
    raise SystemExit('Unknown summary shape')
PY

Verify: http_req_duration, http_req_failed, and http_reqs appear. If they do not, check whether the script made HTTP requests and whether you accidentally opened a real-time JSON stream. Avoid substituting zero for a missing metric; zero latency or errors would look excellent while concealing an incomplete export.

The duration values in the k6 summary are milliseconds, while a rate value is a fraction. For example, a failure rate of 0.02 means 2% failed requests. Keep raw values in JSON and convert them for display only. p(95) describes the 95th percentile of the observed request durations, not a guarantee about future traffic. With only five requests, it is especially unstable.

Checks and thresholds answer different questions. A check records whether each observed response met a condition. A threshold evaluates aggregate metrics and determines the run outcome. For the test here, both check rate and HTTP failure rate must meet their respective rules. If you later add several operations, use stable operation tags and consider operation-specific thresholds, because one fast endpoint can dominate a global percentile.

Step 5: Produce a Comparable Candidate

Run the test again against the same local server without editing the script. Save it under another name. This produces a second observation under nominally identical demand.

k6 run --summary-export candidate.json script.js
python3 -m json.tool candidate.json > /dev/null

Verify: both commands exit successfully and k6 reports five requests and five iterations again. Confirm both artifacts are present:

python3 - <<'PY'
from pathlib import Path
for name in ('baseline.json', 'candidate.json'):
    path = Path(name)
    assert path.is_file() and path.stat().st_size > 0, name
    print(name, path.stat().st_size, 'bytes')
PY

Do not expect identical p95 values. File cache state, laptop scheduling, and the tiny sample count can change timing. The goal is to validate the comparison workflow. For a real baseline, run the same application route mix at the same VUs or arrival rate, iteration duration, environment, and test data. Repeat tests to estimate normal variation. Archive the application revision, script revision, k6 version, target, time, and workload settings alongside each export.

If you change VUs, add an endpoint, or switch from a closed-user executor to an arrival-rate executor, name that a new workload series. It is not a candidate for the old baseline. The k6 scenarios and executors guide explains why configured users and delivered request rate are different quantities. The k6 load testing tutorial covers building an approved workload before collecting comparisons.

Step 6: Compare k6 Summary Export JSON Files With a Strict Parser

Save this as compare.py. It understands the newer array and the older object form, but refuses to compare files of different shapes. It reads three metrics: request count, p95 HTTP duration, and HTTP failure rate. It fails when a required value is absent or nonnumeric instead of treating a broken artifact as an improvement.

import json
import math
import sys
from pathlib import Path


def load_summary(filename):
    with Path(filename).open(encoding='utf-8') as file:
        data = json.load(file)
    results = data.get('results')
    if isinstance(results, dict) and isinstance(results.get('metrics'), list):
        shape = 'machine-readable'
        metrics = {m['name']: m['values'] for m in results['metrics']}
    elif isinstance(data.get('metrics'), dict):
        shape = 'legacy'
        metrics = {name: m.get('values', m)
                   for name, m in data['metrics'].items()}
    else:
        raise ValueError(f'{filename}: unknown summary schema')

    def metric(name, *fields):
        try:
            values = metrics[name]
            value = next(values[field] for field in fields if field in values)
            result = float(value)
        except (KeyError, StopIteration, TypeError, ValueError) as error:
            raise ValueError(f'{filename}: missing {name}.{fields}') from error
        if not math.isfinite(result):
            raise ValueError(f'{filename}: nonfinite {name}.{fields}')
        return result

    requests = metric('http_reqs', 'count', 'value')
    if requests <= 0:
        raise ValueError(f'{filename}: no HTTP requests')
    return {
        'shape': shape,
        'requests': requests,
        'p95_ms': metric('http_req_duration', 'p(95)'),
        'failure_rate': metric('http_req_failed', 'rate'),
    }


def display(label, before, after, unit):
    delta = after - before
    relative = f'{delta / before * 100:+.1f}%' if before else 'n/a'
    print(f'{label}: {before:.3f}{unit} -> {after:.3f}{unit} '
          f'({delta:+.3f}{unit}, {relative})')


if len(sys.argv) != 3:
    raise SystemExit('Usage: python3 compare.py baseline.json candidate.json')
try:
    baseline = load_summary(sys.argv[1])
    candidate = load_summary(sys.argv[2])
    if baseline['shape'] != candidate['shape']:
        raise ValueError('Summary schemas differ; rerun with one k6 release')
    print('Schema:', baseline['shape'])
    display('Requests', baseline['requests'], candidate['requests'], '')
    display('HTTP p95', baseline['p95_ms'], candidate['p95_ms'], ' ms')
    display('Failure rate', baseline['failure_rate'] * 100,
            candidate['failure_rate'] * 100, '%')
except (OSError, json.JSONDecodeError, ValueError) as error:
    raise SystemExit(str(error))
python3 compare.py baseline.json candidate.json

Verify: the output names the schema and reports both values for each metric. Request counts should be five. The newer summary schema defines http_reqs.values.count; some older exports use count or value, which the parser accepts. If your older file has another field, inspect it and adapt the parser deliberately. Never add an unverified fallback that silently converts absent data to zero.

The display for Failure rate multiplies the fraction by 100. Its raw delta is therefore in percentage points, while the final parenthesized value is a relative percent change. A change from 1% to 2% is one percentage point and 100% relative growth. State which one you mean in a review. An increase in p95 with a stable request count and low failure rate is a useful investigation lead, not yet proof of a code regression.

For a real review, add a small comparison record next to the raw files. Include the exact CLI command, the target build identifier, the test data set, and whether a cache warm-up was performed. If an arrival-rate scenario was used, record both the configured arrival rate and dropped_iterations; otherwise the intended load may differ from the work actually executed. If a closed-user scenario was used, response time affects how quickly each user starts the next iteration, so equal VU counts do not guarantee equal throughput. Mark a run inconclusive when the generator is saturated or the target deploy changed during measurement. A good review can explain the observation without implying more precision than the experiment provides. Keep the raw exports because a later analyst may need to recalculate a display or inspect a threshold that this simple comparator does not print.

Step 7: Use the Export in CI Without Losing Failures

Each k6 run has two separate outputs: a file and a process status. A pipeline must preserve both. A failed threshold can still leave an artifact that a later json.tool command parses successfully. Configure your job to stop on a failed k6 command and to upload the artifact in cleanup. The following shell block is a small local CI demonstration; it assumes k6 and Python are already installed on the runner.

set -e
python3 -m http.server 8000 --bind 127.0.0.1 >/tmp/k6-local-server.log 2>&1 &
server_pid=$!
trap 'kill "$server_pid"' EXIT
python3 - <<'PY'
import time
from urllib.request import urlopen
for attempt in range(30):
    try:
        with urlopen('http://127.0.0.1:8000/', timeout=1) as response:
            assert response.status == 200
        break
    except OSError:
        time.sleep(0.1)
else:
    raise SystemExit('local server did not start')
PY
k6 run --summary-export candidate.json script.js
python3 -m json.tool candidate.json > /dev/null
python3 compare.py baseline.json candidate.json

Verify: all commands finish, the comparison prints its three metrics, and the job can archive candidate.json. For a real pipeline, retrieve a baseline from a known successful run and use an authorized staging target. Do not use the local Python server as a capacity gate. A shell's set -e handles this simple sequence, but CI cleanup or artifact upload needs its own always-run setting so failed runs remain inspectable.

A same-run threshold answers whether the result met a specified objective. A baseline comparison answers whether it changed. If you want the comparator to fail a release, first define a rule from repeated runs and acceptable variance. A hypothetical 20% p95 increase may be alarming for a stable high-volume test and meaningless for a five-request laptop test. Also compare delivered work: a faster p95 is not a gain if the candidate handled fewer requests or returned many fast errors. Correlate with CPU, queue, and error evidence in the Grafana dashboards for test metrics guide.

When to Use handleSummary() Instead

For a new integration, Grafana recommends handleSummary() because you control the destination and serialized content. It runs after k6 aggregates results and returns a map from output names to strings or ArrayBuffer values. The separate script below shows the API without changing script.js or the comparator:

import http from 'k6/http';
import { check } from 'k6';

export const options = { vus: 1, iterations: 5 };

export default function () {
  const response = http.get('http://127.0.0.1:8000/');
  check(response, { 'status is 200': (r) => r.status === 200 });
}

export function handleSummary(data) {
  return { 'custom-summary.json': JSON.stringify(data, null, 2) };
}

Save it as custom.js, then run:

k6 run custom.js
python3 -m json.tool custom-summary.json > /dev/null

Verify: the file is valid JSON. Inspect its top-level keys before parsing, because the argument's schema can evolve with k6. This handleSummary() returns only a file, so it replaces the usual terminal summary. Add a stdout destination if your team requires console reporting. The custom summary documentation describes output destinations and data shape.

A custom report lets you choose and version a smaller contract, such as application revision, workload ID, request count, p95, and failed-threshold status. That design takes maintenance: the script and consumer must agree on names, units, and missing-field behavior. Keep the existing --summary-export pipeline until consumers have migrated and test fixtures prove the replacement works.

Troubleshooting

  • Problem: the JSON file is missing. Confirm the output directory exists and is writable, the run reached summary generation, and --summary-mode=disabled is not set. That mode disables summary export too. Print the absolute path in CI when working directories can change.
  • Problem: json.tool reports extra data. You probably supplied --out json=points.json. That is a line-delimited stream, not one summary document. Use --summary-export for this workflow or parse each point separately.
  • Problem: p(95) is absent. Check that HTTP requests occurred and summaryTrendStats includes p(95). Inspect the schema before selecting a field path. A missing percentile is an invalid comparison input, not zero milliseconds.
  • Problem: a candidate appears faster while errors rose. Read http_req_failed and checks before latency. Fast error responses can make global p95 look better. Compare successful operation-specific traffic under the same delivered load.
  • Problem: CI is green after a failed threshold. Preserve the k6 run exit status. A later successful JSON parse can mask it in a shell without fail-fast behavior. Upload the artifact even when the job fails so the result is diagnosable.
  • Problem: baseline and candidate request counts differ. Inspect scenario settings, test duration, target availability, and dropped iterations. Do not claim an application regression until you know why delivered demand changed.

Interview Questions and Answers

The six model answers in interviewQnA below cover export format, threshold status, comparability, missing fields, error-driven latency changes, and migration to handleSummary(). Practice explaining your actual artifact and workload rather than memorizing the flag. The performance testing interview questions guide covers wider performance scenarios.

Common Mistakes

  • Overwriting a baseline with a new summary.json in the same directory. Give artifacts immutable run identifiers and retain the previous successful one.
  • Comparing different scripts or load profiles as if they were the same experiment. Record a workload ID and start a new baseline series after changing demand.
  • Treating checks as process-failing assertions without thresholds. Add checks criteria when correctness must gate CI.
  • Reading the printed, rounded console p95 instead of the raw JSON value. Use the machine-readable artifact and state its units.
  • Inferring improvement from p95 alone. Include failures, delivered requests, checks, and server observations.
  • Calling five local requests a performance benchmark. Use controlled repeated runs and a representative target before making release decisions.

Where To Go Next

Replace the local target with a stable endpoint in a system your team owns. Define one operation and a realistic load profile, run it repeatedly, and keep the exact script, application revision, k6 release, environment, and threshold outcome with every export. Add operation-specific metrics when multiple routes enter the test. The k6 performance engineering guide expands workload design and diagnosis, while performance testing with k6 scripts gives more script patterns.

If you need a durable machine-readable contract, prototype handleSummary() and validate it against the current Grafana summary schema. If you already use --summary-export, keep its behavior explicit during migration. A defensible comparison states which metric changed, by how much, under what delivered demand, and whether both runs still passed their correctness and performance thresholds.

Interview Questions and Answers

What would you validate before comparing two k6 summary exports?

I would check that both runs used the same script revision, route mix, target environment, load profile, summary schema, and trend statistics. I would compare delivered requests and failure rates before latency. If those differ materially, I would investigate test conditions before attributing the p95 change to the application.

How is --summary-export different from --out json?

The summary export is one aggregate end-of-test document. The JSON output backend writes line-delimited points and metric metadata throughout the run. I use the summary for compact comparison and point data when I need a time window or tags for diagnosis.

Why preserve the k6 process exit code when archiving summary.json?

A JSON artifact can be produced even when a threshold fails. The file proves that results were written, not that the performance objective passed. I keep the exit status and threshold details with the artifact so a later parser cannot turn a failed run into a green CI job.

How would you handle a missing p95 value?

I would stop the comparison and inspect the file schema, HTTP request count, and `summaryTrendStats`. A missing percentile can mean no samples or an incompatible export shape. I would not substitute zero, because that falsely looks like a large improvement.

Why can a faster p95 accompany a worse release?

Error responses can complete much faster than successful requests, and a failed generator can deliver less work. I first check `http_req_failed`, functional checks, delivered iterations, and request counts. Then I compare successful operation-specific latency under matched demand.

When would you choose handleSummary() instead of --summary-export?

I would use `handleSummary()` when a new integration needs a stable, deliberate output contract or multiple formats. It lets the script control the destination and serialization. I would version the custom format and test its parser, because the summary data supplied by k6 can evolve.

Frequently Asked Questions

What does k6 --summary-export write?

It writes one JSON file of end-of-test aggregate results, including metrics, checks, and threshold information. It does not contain every request sample. Use `--out json=...` when you need granular points over time.

How do I export a k6 summary to JSON?

Run `k6 run --summary-export summary.json script.js` and confirm the command finishes. Validate the file with `python3 -m json.tool summary.json`. Preserve the k6 exit code because a file can exist after a failed threshold.

Is k6 --summary-export deprecated?

Grafana currently documents the flag as available but discouraged for new integrations. Its docs recommend `handleSummary()` for flexible, controlled exports. Check your installed k6 release before changing an existing pipeline.

Why does my summary JSON have a different shape from an example?

k6 has legacy and newer machine-readable summary formats. Inspect the top-level keys and verify the format against the official schema before writing field paths. A parser for one format should not silently process another.

Can I compare two k6 JSON summaries by p95 alone?

No. Confirm comparable scripts, routes, load profiles, environments, and request counts, then inspect failures and checks. A lower p95 can result from fast error responses or less delivered work.

Does a failed check cause a nonzero k6 exit code?

A check records pass or fail samples but does not alone fail the process. Add a threshold such as `checks: ['rate==1']` to make failed checks affect the run outcome. Keep that exit code when exporting JSON in CI.

What is the difference between --summary-export and --out json?

`--summary-export` writes one aggregate end-of-test JSON document. `--out json=points.json` streams line-delimited metric declarations and sample points while the test runs. Their parsers and use cases are different.

Related Guides