What It Takes to Run an Assessment for 65,000 Candidates at Once
· 7 min read
Contents
Writing a good question paper takes time and money. If candidates take a test in batches, the company has two bad choices. Reuse the paper, and questions leak from one batch to the next. Write several papers, and one may turn out easier than another.
The fairest option is to give everyone the same paper at the same time. That turns a fairness requirement into an engineering problem: at the start time, tens of thousands of candidates log in, check their devices, verify their identity and click Start Assessment within a few minutes of each other.
This is not the same as serving 65,000 page views. Each candidate is in a high-stakes session with state. We have to record exactly when they started, deliver the right paper, and capture every answer, every change of answer, the time spent on each question and every proctoring signal. Losing even one answer we have acknowledged is not acceptable.
Three systems, not one
The most important decision was to stop treating an assessment as one big workflow. We split it into three systems, each with a very different load:
- Publishing. The company writes questions, sets up sections and randomisation, and schedules the test. This data is relational and edited often, so it lives in MySQL.
- Attempting. What candidates use during the test: starting, loading questions, saving answers, proctoring and submitting.
- Evaluation. Starts after submission: scoring, running code, AI evaluation and reports.
Because they are separate, a spike in one does not slow down the others.
Take question delivery off the database
If 65,000 candidates clicked Start and each request built a paper by reading questions from MySQL, the database would get a huge read spike at the worst possible moment.
So papers are prepared before the test begins. Each candidate’s variant is assigned in advance, and the papers are encrypted and served from a CDN, infrastructure built for exactly this kind of high-volume delivery. When the test starts, nothing has to be queried or assembled, and no paper can be opened before its start time.
The best database query during a traffic spike is the one you removed beforehand.
The start is the riskiest moment
Even with papers out of the way, starting is still the hardest part. For each candidate the system may need to:
- log them in
- create the attempt and record the exact start time
- verify the secure exam environment
- confirm their identity
- set up the candidate’s and the assessment’s state
Normal SaaS traffic spreads out over the day. Here, the whole point is that everyone starts together. So we cut down the shared infrastructure each request touches.
Authentication is a good example. Checking every API request against a token in Redis would make Redis a dependency for almost every action a candidate takes. Instead, Redis is used only when a candidate first logs in. After that, every request carries a signed JWT, which the API servers check with local compute: no database lookup and no network call. That makes authentication easy to scale out.
Never lose an answer
The answer is the most important data in the system, and many things can get in its way. A candidate’s internet may drop for a few minutes. A server may be scaling. A request may time out after the candidate has moved on.
So the server is never the only copy. Every answer is written first to IndexedDB in the candidate’s browser. A service worker picks up pending answers and sends them to the server. If a request fails, the answer stays in the browser and the service worker retries with exponential backoff. When the connection comes back, the queued answers carry on to the server.
On the server, answers go into Cassandra, which suits a write-heavy load and scales out.
We keep more than the final choice. If a candidate changes an answer from A to B to C, we record each change and the time between them. An attempt is a stream of events, not just a final answer sheet.
That is what durability means here: retries are safe, state can be rebuilt after an interruption, and an older event never overwrites a newer one.
Keep evaluation away from live attempts
Evaluation is expensive, but it does not have to happen right away. A candidate should not wait for scoring, code runs, AI evaluation or reports before we confirm their submission.
After the test, answers are copied from Cassandra into a separate database. Evaluation workers read from a replica and work through their own queues. Results go to a secondary store and are moved back into the main MySQL systems at a controlled pace.
If the queues back up, reports arrive later, but candidates still taking the test are not affected. Submission accepted and evaluation finished are two different events, and the system treats them that way.
Proctoring at this scale
Proctoring adds another constant stream of data: signals from the secure exam environment, from the candidate’s device and, for stricter tests, from additional monitoring.
Sending every raw signal to central servers would need enormous bandwidth and compute. So most of the work happens before the data reaches our core systems, and only what matters is kept. Proctoring produces evidence and flags without competing with answer saves for the infrastructure that matters most.
The lesson
At this scale, adding more app servers is the easy part. The hard part is protecting what is shared and stateful: databases, auth stores, write pipelines, queues and evaluation.
There is no single database or cloud service that solves this. The answer is to understand how each part of the workflow behaves and design each one on its own terms:
| Part | What it needs |
|---|---|
| Publishing questions | Relational consistency and easy editing |
| Delivering questions | Prepared ahead of time, encrypted, served from a CDN |
| Capturing answers | A local copy, safe retries and writes that scale out |
| Authentication | No network calls it can avoid |
| Evaluation | Its own workers, queues and databases |
| Proctoring | Early filtering and controlled ingestion |
We were not just scaling web traffic. We were looking after thousands of career-defining moments happening at once. Success was not only that the APIs stayed up. It was that every candidate could trust they got the right paper, every answer was kept, and the infrastructure had no say in their result.