Resource library

QA How-To

k6 setup and teardown Tutorial: Share an Auth Token Across VUs

Use this k6 setup and teardown tutorial to log in once, share an auth token across VUs, isolate request metrics, and verify cleanup with runnable code.

19 min read | 2,903 words

TL;DR

Authenticate once in `setup()`, return `{ token, sessionId }`, use that data in each VU's protected request, and revoke the token in `teardown()`. Verify one login and one revocation with server-side counters; give sessions an independent expiry because interrupted tests may skip cleanup.

Key Takeaways

  • Return a compact JSON object from setup() to give each VU the same initial authentication data.
  • Verify the shared session by comparing a protected response's session ID with setup data.
  • Tag business requests so login and revocation do not affect their latency thresholds.
  • Use teardown() for normal cleanup and server-side expiry for interrupted runs.
  • Move login into VU code when the workload must represent distinct users or authentication traffic.
  • A setup exception prevents teardown(), so handle partially created resources inside setup.

A k6 setup and teardown tutorial is most useful when you can prove the lifecycle behavior, not just print a token. In this guide, setup() authenticates once, returns a small JSON object to four VUs, and teardown() revokes that same session after the load work ends. A local API exposes safe counters so you can verify one login and one revocation without logging a credential.

The pattern fits a service account or a shared test session whose authorization scope is identical for every VU. It does not model four distinct users. If your system enforces a single active session per user, refreshes tokens under load, or needs per-user authorization, give each VU its own identity instead. See k6 scenarios and executors for choosing the appropriate load shape.

What You Will Build

  • A local Python HTTP API with login, protected profile, session revocation, and counter endpoints.
  • A k6 script that returns { token, sessionId } from setup() and reads that data in every VU.
  • Checks and thresholds scoped to the protected business request, so login and cleanup traffic do not distort its latency result.
  • A verification run with four VUs and a failure run that confirms cleanup still happens after a VU assertion fails.

You will create the API and k6 script in a scratch directory. The local credentials are deliberately disposable. In a real environment, supply credentials through your secret manager and avoid including them in a repository or test output.

Prerequisites

Use Grafana k6 2.3.0, the version listed in the current release notes, and Python 3.12.7, the Python version used for the static checks of this example. Install k6 with the official platform instructions; the script needs only Python standard-library modules. If your installed versions differ, record the exact output below and check the matching k6 documentation before running. The commands below assume a POSIX shell and curl. On Windows, run them in a compatible shell or translate the file creation commands to PowerShell.

k6 version
python3 --version
curl --version
mkdir -p k6-auth-demo
cd k6-auth-demo

Verify: All three version commands should succeed, and pwd should end in k6-auth-demo. If k6 is missing, finish its platform-specific installation before starting the server. Check the installed binary's version against the current k6 documentation because CLI availability can change across releases. No Docker image tag or package version is assumed here.

The API listens only on 127.0.0.1:8000. Reserve that port or set a different port consistently in both files. The exercise uses real HTTP calls and real k6 lifecycle hooks, but its load target is a local fixture; latency numbers from it are useful for verifying the script, not for judging a production service.

Step 1: Start a Local Authentication Fixture

Create mock_api.py with three application endpoints. The server keeps sessions in memory, records successful logins and revocations, and never returns the bearer token through its debug endpoint. A lock protects the shared dictionary because ThreadingHTTPServer handles concurrent VU requests on separate threads. The optional FAIL_PROFILE=1 switch gives you a controlled failure for Step 6.

cat > mock_api.py <<'PYFILE'
import json
import os
import secrets
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from threading import Lock
from urllib.parse import urlsplit

sessions = {}
counters = {"login_calls": 0, "revoke_calls": 0}
state_lock = Lock()


class Handler(BaseHTTPRequestHandler):
    def send_json(self, status, payload):
        data = json.dumps(payload).encode("utf-8")
        self.send_response(status)
        self.send_header("Content-Type", "application/json")
        self.send_header("Content-Length", str(len(data)))
        self.end_headers()
        self.wfile.write(data)

    def bearer_token(self):
        header = self.headers.get("Authorization", "")
        return header[7:] if header.startswith("Bearer ") else ""

    def do_POST(self):
        if urlsplit(self.path).path != "/auth/login":
            self.send_json(404, {"error": "not found"})
            return
        try:
            length = int(self.headers.get("Content-Length", "0"))
            credentials = json.loads(self.rfile.read(length))
        except (ValueError, json.JSONDecodeError):
            self.send_json(400, {"error": "invalid JSON"})
            return
        if (credentials.get("username") != os.getenv("TEST_USER", "load-user")
                or credentials.get("password") != os.getenv("TEST_PASSWORD", "local-only-secret")):
            self.send_json(401, {"error": "bad credentials"})
            return
        token = secrets.token_urlsafe(32)
        session_id = secrets.token_hex(8)
        with state_lock:
            sessions[token] = session_id
            counters["login_calls"] += 1
        self.send_json(200, {"token": token, "sessionId": session_id})

    def do_GET(self):
        path = urlsplit(self.path).path
        if path == "/debug/state":
            with state_lock:
                result = {**counters, "active_sessions": len(sessions)}
            self.send_json(200, result)
            return
        if path != "/api/profile":
            self.send_json(404, {"error": "not found"})
            return
        with state_lock:
            session_id = sessions.get(self.bearer_token())
        if session_id is None:
            self.send_json(401, {"error": "unauthorized"})
        elif os.getenv("FAIL_PROFILE") == "1":
            self.send_json(503, {"error": "planned profile failure"})
        else:
            self.send_json(200, {"sessionId": session_id, "name": "Load User"})

    def do_DELETE(self):
        if urlsplit(self.path).path != "/auth/session":
            self.send_json(404, {"error": "not found"})
            return
        with state_lock:
            existed = sessions.pop(self.bearer_token(), None)
            if existed is not None:
                counters["revoke_calls"] += 1
        if existed is None:
            self.send_json(401, {"error": "unauthorized"})
            return
        self.send_response(204)
        self.end_headers()


ThreadingHTTPServer(("127.0.0.1", 8000), Handler).serve_forever()
PYFILE
python3 mock_api.py

Leave the server running in that terminal. In a second terminal, enter the same directory and query the counter endpoint. The first response should show login_calls: 0, revoke_calls: 0, and active_sessions: 0; the exact JSON spacing is irrelevant.

cd k6-auth-demo
curl -fsS http://127.0.0.1:8000/debug/state

Verify: Expect HTTP 200 and zero counters. If the connection fails, check that the Python process is still running. If port 8000 is occupied, stop the other service or change the address in both the Python bind call and the k6 BASE_URL used below. The debug endpoint is intentionally for this local fixture; do not expose a similar endpoint on a production auth service.

Step 2: Write the k6 Setup and Teardown Tutorial Script

Create lifecycle.js. k6 evaluates its top-level code in the init context, where imports and options belong. Network calls belong in lifecycle functions: setup() logs in, the default function exercises the protected endpoint, and teardown() revokes the login. The setup return value is JSON data copied to each VU and separately supplied to teardown. It is not a mutable object shared live between VUs.

cat > lifecycle.js <<'JSFILE'
import http from 'k6/http';
import { check, sleep } from 'k6';
import { Rate } from 'k6/metrics';

const baseUrl = __ENV.BASE_URL || 'http://127.0.0.1:8000';
const username = __ENV.TEST_USER || 'load-user';
const password = __ENV.TEST_PASSWORD || 'local-only-secret';
const profileChecks = new Rate('profile_checks');

export const options = {
  scenarios: {
    shared_session: {
      executor: 'per-vu-iterations',
      vus: 4,
      iterations: 2,
      maxDuration: '30s',
    },
  },
  thresholds: {
    profile_checks: ['rate==1'],
    'http_req_failed{kind:business}': ['rate<0.01'],
    'http_req_duration{kind:business}': ['p(95)<1000'],
  },
};

export function setup() {
  const response = http.post(
    `${baseUrl}/auth/login`,
    JSON.stringify({ username, password }),
    {
      headers: { 'Content-Type': 'application/json' },
      tags: { kind: 'setup' },
    }
  );
  if (response.status !== 200) {
    throw new Error(`Login failed with HTTP ${response.status}`);
  }
  const payload = response.json();
  if (!payload || typeof payload.token !== 'string' ||
      typeof payload.sessionId !== 'string') {
    throw new Error('Login response lacks token or sessionId');
  }
  console.log(`setup: created session ${payload.sessionId}`);
  return { token: payload.token, sessionId: payload.sessionId };
}

export default function (data) {
  const response = http.get(`${baseUrl}/api/profile`, {
    headers: { Authorization: `Bearer ${data.token}` },
    tags: { kind: 'business' },
  });
  let actualSessionId;
  if (response.status === 200) {
    try {
      actualSessionId = response.json('sessionId');
    } catch (error) {
      actualSessionId = undefined;
    }
  }
  const passed = check(response, {
    'profile returns 200': (r) => r.status === 200,
    'profile uses setup session': () => actualSessionId === data.sessionId,
  });
  profileChecks.add(passed);
  console.log(`VU ${__VU}: session ${data.sessionId}, passed ${passed}`);
  sleep(0.2);
}

export function teardown(data) {
  const response = http.del(`${baseUrl}/auth/session`, null, {
    headers: { Authorization: `Bearer ${data.token}` },
    tags: { kind: 'teardown' },
  });
  if (response.status !== 204) {
    throw new Error(`Revocation failed with HTTP ${response.status}`);
  }
  console.log(`teardown: revoked session ${data.sessionId}`);
}
JSFILE
k6 run lifecycle.js

Verify: The output should contain one setup: created session ... line, eight VU ... lines, and one teardown: revoked session ... line. The summary should show eight completed iterations because four VUs run two iterations each. The profile_checks threshold should pass. The script logs a session ID, never the bearer token or password; session IDs can still be sensitive in some systems, so remove that diagnostic logging when adapting the example.

The login response is validated before setup() returns. If it is not HTTP 200 or misses either field, the test fails before any VU starts. k6 does not call teardown() when setup() aborts, so a real setup workflow that allocates resources before the login must clean them up within setup's own error handling. The fixture allocates its session only after it validates credentials, which keeps this example's failed-login path simple.

Step 3: Verify the Same Token Reaches Every VU

Read the second terminal's k6 output. Each of the four VU numbers should print the same sessionId, and each line should say passed true. That is stronger evidence than a successful HTTP status alone: the protected response contains the server's session ID, which the VU compares with the ID returned by the single setup login. A different token might still yield HTTP 200, but it would not match this exact session.

Run a clean comparison by restarting the Python server, which resets its in-memory counters, and then running k6 once. Restarting also clears any session accidentally left behind by an earlier interrupted run. Execute the following in the k6 terminal after the new server is listening:

k6 run lifecycle.js
curl -fsS http://127.0.0.1:8000/debug/state

Verify: The state should be {"login_calls": 1, "revoke_calls": 1, "active_sessions": 0} after this clean run. Four VUs did not trigger four logins; they received copies of the one setup result. If you did not restart the fixture, compare differences instead: each completed k6 run should add exactly one to each call counter and leave zero active sessions.

The data argument to default(data) is stable for that VU, but changing it inside one VU does not modify another VU's copy or the object later passed to teardown(data). This distinction matters for workflows that create a list of records under load: teardown cannot infer IDs created by VUs just by reading setup data. If cleanup needs those IDs, record them through an external service designed for that purpose, or pre-create the resources in setup and return their IDs. Avoid returning a large fixture from setup because k6 must serialize and distribute that data to VUs.

Approach Login volume in this example Identity visible to four VUs Use when
One login in setup() One request per test run One shared service account Measuring a protected operation under a common identity
Login inside default() Up to eight requests across eight iterations Depends on the credentials assigned in VU code Login is part of the measured journey
One login per VU with assigned accounts Four requests if each VU caches its own token Four distinct accounts Authorization, quotas, or session isolation matter

The table's counts describe this four-VU, two-iteration exercise, not a general k6 default. A per-VU token cache must live in that VU's own runtime state; returning one token from setup cannot produce four independent sessions. Decide which requests belong in your workload before comparing results between these approaches.

Step 4: Separate Authentication Traffic From Business Metrics

Inspect the kind tags in lifecycle.js. The login request uses kind: setup, profile requests use kind: business, and revocation uses kind: teardown. k6 still counts all HTTP requests in the global totals. The thresholds select only business samples so one unusually slow login does not alter the protected endpoint's p95, and a deliberate failure of the cleanup API does not masquerade as a business request failure. See k6 thresholds and checks for a deeper discussion of checks versus pass/fail thresholds.

The custom profile_checks rate turns both profile assertions into one iteration-level success value. check() reports its named assertions, but checks by themselves do not fail the k6 process. The rate threshold makes any failed profile iteration a failed test. The http_req_failed{kind:business} threshold catches transport and HTTP failures; the duration threshold protects the endpoint's p95. The 1000 millisecond limit is illustrative for the local fixture, not an SLO for your application.

Run a summary with the same script and inspect the threshold lines:

k6 run --summary-mode=full lifecycle.js

Verify: Look for profile_checks, http_req_failed{kind:business}, and http_req_duration{kind:business} in the output, each passing on a healthy local run. If your installed k6 release does not recognize --summary-mode=full, run k6 run lifecycle.js and use its normal end summary; the threshold keys remain the same. Confirm that the request count includes login and revocation in addition to the eight protected calls. For a meaningful performance report, compare the tagged profile latency to a separately defined service objective, then choose a workload in the k6 load testing tutorial that reflects expected arrival patterns.

A shared token deliberately removes login traffic from the repeated VU path. That is correct when you want to measure a protected endpoint with an already authenticated service identity. It underestimates authentication load if real users log in during the journey. For an end-to-end user journey, move login into the VU flow, assign independent accounts, and measure authentication as its own operation. State explicitly which model you used when presenting results.

Step 5: Prove Teardown Revokes the Session

The successful run ends with a DELETE request to /auth/session. The Python server removes the token from its active map, increments revoke_calls, and returns HTTP 204. The teardown function throws if the server returns any other status, making cleanup failure visible. It receives the original setup result even though the VUs have executed their own copies of that data.

Run the same test once more and check the counters directly. The first command records the current state, and the second run should increase both call counters by one while keeping active sessions at zero:

curl -fsS http://127.0.0.1:8000/debug/state
k6 run lifecycle.js
curl -fsS http://127.0.0.1:8000/debug/state

Verify: Compare the two JSON objects: login_calls increases by 1, revoke_calls increases by 1, and active_sessions ends at 0. A missing teardown line plus one active session suggests the run was interrupted, setup failed after allocating a resource, or the cleanup request failed. Investigate the k6 exit code and server log before rerunning. The counters provide direct evidence of lifecycle behavior without exposing a token.

Do not depend on teardown as the sole safeguard for scarce resources. A killed process, machine crash, or setup exception can bypass it. Give test accounts and sessions bounded lifetime server-side, scope them to a dedicated environment, and have an independent cleanup path for orphaned data. If your API supports idempotent revocation, prefer it; a retry after an uncertain network response should be safe. This fixture returns 401 when a token has already been revoked so the tutorial can surface a duplicate delete clearly.

Step 6: Exercise a VU Failure and Watch Cleanup

A failing protected request should cause the run to fail through profile_checks while allowing the ordinary test lifecycle to reach teardown. Stop the Python process with Ctrl+C. Restart it from the same directory with FAIL_PROFILE=1, which makes authenticated profile requests return HTTP 503 but leaves login and revocation working. Then run k6 from the other terminal:

FAIL_PROFILE=1 python3 mock_api.py
k6 run lifecycle.js
curl -fsS http://127.0.0.1:8000/debug/state

Verify: profile returns 200 and profile_checks fail; the k6 process exits nonzero because its thresholds fail. You should still see teardown: revoked session ..., and the debug state should show one login, one revocation, and zero active sessions on this freshly restarted server. The script does not attempt to parse JSON from the 503 response as a successful profile, so it records a failed check without throwing inside the VU.

This demonstration covers a normal test execution that finishes with failed checks. It does not promise cleanup after every possible interruption. In particular, Grafana's k6 lifecycle documentation states that teardown does not run when setup ends abnormally. Restore the healthy fixture after the experiment by stopping the failure-mode server and starting python3 mock_api.py again. Keep a single server process bound to port 8000 at a time.

Step 7: Apply the k6 Setup and Teardown Tutorial to Your Real API

Replace the local URLs and response field names with your application's login, protected operation, and logout contract. Preserve the lifecycle signatures and explicit validation. If your login returns access_token instead of token, map it in setup before returning the compact object. If logout requires a refresh token or a session ID, return that additional field from setup and pass it to teardown. Keep BASE_URL, TEST_USER, and TEST_PASSWORD external to the script. For example, after obtaining a test account from your secret manager, run:

export BASE_URL='https://staging.example.test'
export TEST_USER='dedicated-load-account'
export TEST_PASSWORD='<value-from-your-secret-manager>'
k6 run lifecycle.js

Verify: In the staging server's auth audit log, find exactly one successful login and one successful revocation for the run. Confirm that each protected response belongs to the expected identity, and review the tagged business thresholds. Replace the example host, endpoints, and field mapping before executing this command; example.test is a placeholder, not a reachable service. Use a test account with the minimum permissions required to read the chosen endpoint.

A token's lifetime must exceed the test duration plus setup and shutdown overhead. A long soak test cannot safely reuse a short-lived access token indefinitely. Choose a per-VU refresh workflow, a test-specific long-lived credential with strict scope, or a bounded test duration according to your security policy. Testing bearer token refresh covers the refresh edge cases. If the API invalidates earlier tokens whenever the same account logs in, concurrent tests may interfere; allocate a separate identity per test run even if the VUs within one run share its token.

Interview Questions and Answers

These short prompts test whether you understand the execution model, not just the syntax. The expanded answers in the article's interview set cover the same implementation decisions.

Q: How often does setup() run? Once per test execution, before VU iterations begin.

Q: Can init code send the login request? No. Put HTTP requests in setup or a VU function.

Q: Does a VU mutate data seen by teardown? No. Each stage receives its own copy of setup data.

Q: Why tag login and profile requests differently? A tagged threshold isolates business latency from authentication and cleanup traffic.

Q: What happens if setup throws? VUs do not start, and teardown is not called.

Q: Why can a shared token misrepresent user load? It removes per-user login traffic and may bypass user-specific authorization behavior.

Troubleshooting

Problem: k6: command not found -> Install k6 through the official platform instructions, then run k6 version in the same shell. A copied command from another machine can point to a binary that is not on your PATH.

Problem: curl cannot connect to port 8000 -> Keep python3 mock_api.py running in its own terminal. If Python prints an address-in-use error, stop the process already bound to the port or change the bind address and BASE_URL together.

Problem: setup reports HTTP 401 -> Compare TEST_USER and TEST_PASSWORD in the k6 environment with the variables used to start the fixture. The local defaults match, but exporting only one side changes the credentials for one process.

Problem: profile returns HTTP 401 after setup succeeds -> Check for an accidentally restarted fixture, which erases its in-memory sessions. On a real service, inspect token audience, scope, expiry, and whether a second login invalidated the first token; JWT authentication testing gives focused checks.

Problem: profile_checks fails while HTTP status is 200 -> Inspect the sessionId field in the protected response. A 200 from a different account or a changed response schema cannot satisfy the identity assertion; do not remove the check just to turn the run green.

Problem: revocation fails or an active session remains -> Read the teardown HTTP status and the fixture's counter state. An interrupted process can skip cleanup; revoke orphaned sessions through a separate administrative path and add server-side expiry before running longer tests.

Where To Go Next

Keep this fixture as a small regression exercise for k6 lifecycle behavior. Then replace the local endpoint with a staging API and review the workload model in the complete k6 performance engineering guide. Use the API performance testing tutorial to turn the protected request into a repeatable measurement plan. For account-specific authorization, build a per-VU login flow and compare its results with this shared-session baseline. Record the load shape, token policy, threshold values, and cleanup evidence alongside every report.

Conclusion

Use setup() for one-time authentication, return a small JSON token object, consume that data in each VU, and use teardown() to revoke the session. The local counter endpoint makes the contract observable: one login, multiple protected requests, one revocation, and no active session after a completed run. Before using the pattern against your application, decide whether a shared identity represents the traffic you need to measure and provide an independent cleanup strategy for interrupted runs.

Interview Questions and Answers

Explain the order of k6 init, setup, VU, and teardown stages.

Init loads modules and defines options without making HTTP calls. Setup runs once before the workload and can return JSON data. The VU function executes for its configured iterations, and teardown runs once after a normal workload to clean up setup resources.

How do you share one authenticated session across k6 virtual users?

Send the login request in setup and return the access token with any safe correlation ID. Accept that object as the argument of the VU function, then attach its token as a Bearer header. Verify the server sees the intended session rather than relying only on HTTP 200.

What data types can setup pass to VUs and teardown?

Use JSON-serializable data such as strings, numbers, arrays, and plain objects. Functions and live client objects cannot be transferred through the setup return value. Keep it small because distributing a large object increases memory and serialization work.

Can one VU change the token that another VU or teardown receives?

No. k6 supplies independent copies of the setup return value. A mutation in one VU remains local to that VU and is not a coordination channel for cleanup.

What happens to teardown when setup throws after creating a resource?

Teardown is not called after an abnormal setup exit. Setup must catch failures around multi-step provisioning and clean up any earlier allocations itself. Server-side expiration provides a further backstop for process crashes.

How would you exclude login time from a protected endpoint latency SLO?

Tag the login request as setup and the protected VU request as business. Apply `http_req_duration{kind:business}` to the endpoint's latency threshold. Keep global metrics visible so the authentication and cleanup traffic are still accounted for separately.

Why might a shared token be the wrong performance model?

It makes all VUs act as one principal, which may share caches, quotas, and server-side session state. Real users may each log in and have separate permission checks. Use per-VU accounts when those behaviors are part of the target workload.

How would you handle an access token that expires during a soak test?

Measure the token lifetime against the run length and expected clock skew first. For long runs, use an explicit refresh workflow or per-VU authentication strategy that matches production behavior. A single setup token with no refresh will eventually turn protected requests into authorization failures.

Frequently Asked Questions

Does k6 setup run once per VU?

No. k6 calls setup once before the VU workload. It serializes the returned JSON data and supplies a copy to each VU and to teardown.

Can I return a bearer token from k6 setup?

Yes. Return it as a string in a small JSON object and read it from the default function's data argument. Avoid printing the token, and use a dedicated account with narrow permissions.

Can teardown read values changed by VUs?

No. VUs receive separate copies of setup data, and their mutations are not merged into teardown's argument. Use an external store or pre-created resource IDs when cleanup depends on work performed by VUs.

Will teardown run if setup fails?

No. An abnormal setup exit prevents the VU stage and teardown. Clean up partial allocations inside setup's error path and use server-side expiry for orphaned resources.

Do setup and teardown HTTP requests affect k6 metrics?

They can appear in aggregate HTTP metrics. Tag requests by lifecycle role and apply business latency thresholds only to tagged VU requests when that is the metric you intend to judge.

Is one shared token realistic for a user load test?

Only when your workload genuinely uses one shared service identity. It omits per-user login volume and can hide account-specific access rules, so use separate credentials when modeling distinct users.

How can I confirm all VUs used the same login?

Have the protected endpoint return a safe session identifier and compare it with the identifier returned by setup. Separately inspect server-side login counts; a successful status alone does not prove token identity.

Related Guides