How we load test a checkout
The number you want is not how many users the system can take. It is what breaks first, at how many, and what to change.
A load test that ends with a single number is a load test that will be misread. "It handled 2,000 users" means nothing without the shape of the traffic, the latency people saw along the way and the first thing that gave out. So a checkout load test, for us, has four parts: a forecast, a script that behaves like a customer, a ramp that finds the edge, and a diagnosis.
Start from the forecast, not the hardware
Before we write a script we ask what the busiest hour will look like: how many people, arriving how fast, doing what. A sale with a countdown behaves differently from a slow Tuesday. The forecast becomes the target: for example p95 under 800 ms and under half a percent of errors at 1,500 concurrent users. If nobody has a forecast, we build one from analytics and the marketing plan, and write down that it is a guess.
Make the script behave like a person
- Browse, search, open a product, add to the cart, pay with a test card. Not one endpoint hammered in a loop.
- Real think time between steps, drawn from a range, so requests do not arrive in lockstep.
- Unique users with unique carts. Caches are flattering when every virtual user buys the same thing.
- A production-like environment: same database size, same third parties stubbed at realistic latency, same CDN rules.
Ramp until something gives
We ramp well past the target, slowly enough to see where latency starts to curve, then hold at the target to find leaks. The interesting number is the knee: the load at which p95 stops growing gently and starts growing fast. In a recent sample that knee was about 1,400 users, under the forecast, which is exactly why the test was run.
Diagnose, then prescribe one thing
A knee always has a cause: a connection pool that is too small, a query with no index, a lock, a third party with a rate cap. We find it with the metrics the client already has, name it, and recommend the change that moves the knee furthest for the least work. Then we run the same test again after the change, so the report ends with a before and an after rather than a hope.
The deliverable is a page with a scenario, a results table, a sentence about where it broke, the fix and the retest. The load test sample on the home page shows the shape.