← All posts
March 10, 2025·8 min read

We load tested every mock API service so you don't have to

We ran k6 load tests against Mockra, Postman Mock Server, WireMock Cloud, and Stoplight Prism at 1000 concurrent virtual users. Here's what we found — including results that surprised us.

You're running a k6 load test with 1000 virtual users. Your application is humming along. Then you notice the bottleneck isn't your application — it's the mock server standing in for the third-party API you depend on.

This is a real failure mode. It's not theoretical. We built Mockra because we lived this: a performance test that told us nothing useful because the mock layer — not the application — was the constraint.

So we decided to test every popular mock API service under real load conditions and publish the results.

Conflict of interest disclosure: We built Mockra. We ran this benchmark. We have an obvious incentive to present favorable results. We've tried to mitigate that by publishing all scripts, all raw data, and all methodology so you can run the same tests yourself and dispute our numbers if they're wrong. The link to do that is at the bottom.


Methodology

We chose the simplest possible test case: a GET /api/v1/users endpoint that returns a JSON array of two users. Complexity is not the variable being tested — throughput is. Every platform was given the same OpenAPI spec and configured to return the same response.

We used k6 as the test runner with this load profile:

stages:
  - duration: 2m, target: 100    # ramp up
  - duration: 2m, target: 500    # mid ramp
  - duration: 2m, target: 1000   # full load
  - duration: 5m, target: 1000   # sustained steady state
  - duration: 1m, target: 0      # ramp down

Platforms tested:

  • Mockra — our own product, edge-deployed on Cloudflare Workers
  • Postman Mock Server — free tier, the most widely-used mock tool in the market
  • WireMock Cloud — Developer (free) tier, the cloud offering of the popular open-source WireMock
  • Stoplight Prism — self-hosted via Docker on the same test machine (loopback)

Test machine: AWS t3.xlarge (4 vCPU, 16 GB RAM) in us-east-1. k6 v0.55.0. All hosted services tested from the same region where available.

What we measured:

  • Sustained requests per second at 1000 VUs
  • P50 / P95 / P99 response time
  • Error rate (non-2xx responses)
  • Cold start behavior (first request after 10 minutes idle)

Results

Mockra

Mockra is built on Cloudflare Workers. Requests are served from the edge point-of-presence (PoP) closest to the caller. There are no cold starts — Workers are always warm at the edge.

Metric Value
Sustained RPS 6,695
P50 latency 7ms
P95 latency 18ms
P99 latency 31ms
Error rate 0.02%
Cold start (10 min idle) 11ms

At 1000 VUs, Mockra sustained 6,695 requests per second with a P95 of 18ms. Both thresholds passed comfortably. The cold start test showed 11ms on the first request after a 10-minute idle period — effectively indistinguishable from steady-state latency. Cloudflare Workers don't spin down between requests; the edge PoP keeps them warm.

The 0.02% error rate (892 errors across 4.8 million requests) is noise — connection resets and the occasional timeout at peak concurrency, not systematic failures.

WireMock Cloud

WireMock Cloud is the hosted offering of WireMock, which runs on the JVM. It performed well compared to Postman and passed both thresholds, but with notably higher latency and one significant surprise: cold start behavior.

Metric Value
Sustained RPS 2,560
P50 latency 135ms
P95 latency 384ms
P99 latency 813ms
Error rate 1.97%
Cold start (10 min idle) 1,840ms

WireMock Cloud handled the load without rate-limiting — it sustained 2,560 RPS with a 1.97% error rate (errors concentrated in the late ramp-up phase). P95 at 384ms is within the 500ms threshold, but only just.

The cold start figure deserves attention: 1,840ms on the first request after 10 minutes idle. This is JVM warm-up behavior. WireMock Cloud appears to unload idle deployments and restart on demand. If your CI/CD pipeline creates a fresh test environment, that first request — the one that kicks off your test — will take nearly 2 seconds. Depending on how your test is configured, this may affect your ramp-up numbers.

Stoplight Prism (self-hosted)

Prism is a different category from the other three: it's a self-hosted tool you run yourself, not a managed cloud service. We tested it via Docker on the same machine as the k6 runner, so the numbers include zero network latency. This is an important caveat — in practice, Prism would be running somewhere on your network, adding 5–50ms of RTT that isn't reflected here.

Metric Value
Sustained RPS 3,034
P50 latency 28ms
P95 latency 199ms
P99 latency 1,244ms
Error rate 10.0%
Cold start N/A (loopback, no idle spindown)

Prism's P95 at 199ms looks competitive, but there are two problems.

First, it's loopback — the latency numbers aren't comparable to the hosted services. Add 20–80ms of realistic network RTT and Prism's P95 climbs to 219–279ms at minimum before any load-related degradation.

Second, Prism hit a 10% error rate. A single Node.js process saturated at around 650 concurrent VUs. Above that threshold, the event loop backed up and requests started timing out. The error rate peaked at 18.2% during the sustained 1000 VU phase. In a real scenario, you'd need to run multiple Prism instances behind a load balancer — which adds operational complexity that the hosted services absorb for you.

Prism is useful for local development and low-concurrency API contract testing. It is not designed for load testing at scale, and shouldn't be evaluated as if it were.

Postman Mock Server

Postman Mock Server is the most widely used tool in this category, included in millions of developer workflows. It is not designed for load testing, and the numbers reflect that.

Metric Value
Sustained RPS 949
P50 latency 287ms
P95 latency 632ms
P99 latency 1,204ms
Error rate 21.7%
Cold start (10 min idle) 487ms

Postman Mock Server began returning HTTP 429 (Too Many Requests) at approximately 312 concurrent VUs. By the time we reached 1000 VUs, 21.7% of requests were being rate-limited. Both thresholds failed.

This is not a criticism of Postman as a product. Postman Mock Server is designed for API development workflows — sharing examples with your team, testing a new endpoint, demonstrating a spec. It is not designed to absorb 1000 simultaneous load test VUs. The rate limits exist because that's not the product's purpose.

If you're using Postman mocks as your load test target, you're not testing your application. You're testing Postman's rate limiter.


Side-by-side comparison

Platform Sustained RPS P95 Latency Error Rate Cold Start Thresholds
Mockra 6,695 18ms 0.02% 11ms ✅ Pass
WireMock Cloud 2,560 384ms 1.97% 1,840ms ✅ Pass
Prism (local)* 3,034 199ms 10.0% N/A ❌ Fail
Postman 949 632ms 21.7% 487ms ❌ Fail

*Prism tested via loopback — latency not comparable to hosted services.


What this means for load testing

If your load test needs more than 300 RPS with sustained concurrency, Postman Mock Server becomes the bottleneck. Your k6 report will show high error rates and P95s that have nothing to do with your application. You'll tune the wrong thing.

If you need more than ~2,500 RPS or care about cold start behavior in CI pipelines, WireMock Cloud will limit you — and that 1.8-second JVM warm-up will pollute your ramp-up metrics.

If you're running Prism locally, it works fine for development. But a single container can't handle more than ~650 concurrent VUs before errors start accumulating. Running load tests against a local process on the same machine also means you're sharing CPU with k6 — another confound.

The practical decision point: if your load test VU count is below 100, any of these tools will work for light testing. Above 300 VUs, you need infrastructure designed for the workload. Cloudflare Workers handle the concurrency natively — there's no scaling decision to make.


Raw data and reproduce it yourself

All k6 scripts, the OpenAPI spec, and the raw results JSON are in the mockra-io/benchmarks GitHub repository.

The README walks through exactly how to set up each platform and run the tests. The whole process takes about 30 minutes per platform once accounts are configured.

If you run the tests and get meaningfully different results, open an issue. We want the numbers to be accurate, even if accuracy doesn't favor us.


Try Mockra

If you've been using Postman mocks or a local Prism instance as your load test target, Mockra is a drop-in replacement with no infrastructure to manage.

Upload your OpenAPI spec, get a mock URL (mock.mockra.io/your-namespace/...), point your k6 script at it. The free tier is 5,000 requests/month — enough to validate the setup before committing.

Or try error.mockra.io — a free endpoint that returns random realistic error responses (401s, 429s, 500s, 503s) drawn from 1,000 pre-generated messages. No account required. Useful for testing your application's error handling paths under load.

Try Mockra free

5,000 requests/month, no credit card required. Upload an OpenAPI spec and get mock endpoints in under 60 seconds.

Start for freeTry error.mockra.io