Impact
Fairness at scale: Maki’s hybrid approach to proctoring
How can we make sure every candidate takes their test under the same conditions?
October 8, 2025
Maxime Legardez-Coquin

At Maki People, our mission has always been to make hiring both fair and efficient through reliable, science-driven assessments. But as more tests move online, a fundamental question arises: how can we make sure every candidate takes their test under the same conditions?
Cheating undermines trust. It damages the integrity of results and can hurt a company’s reputation. Yet being overly strict can be just as harmful—penalizing honest candidates and creating frustration. The real challenge is finding the balance between accuracy and fairness.
That’s where Maki’s hybrid proctoring technology comes in. It’s a system that combines the precision of computer vision with the reasoning power of large language models (LLMs) to spot genuine cheating while avoiding false positives. The result is a proctoring method that is accurate, explainable, and scalable, giving companies the confidence that their hiring decisions are based on reliable, unbiased results.
The hidden complexity of online cheating
Cheating isn’t always obvious. It can range from someone helping off-camera to a candidate using earbuds, a second screen, or even a phone placed just outside the frame. Sometimes, the signs are subtle - a reflection mistaken for another person, or a photo on the wall misread as a face. The challenge isn’t just detecting these anomalies; it’s interpreting them intelligently.
A fair system needs to understand context. That’s why our approach combines two types of AI models: one that sees and one that reasons.
Figure 01 · The trade-off
There are two ways to get proctoring wrong
Every detection system sits somewhere on this dial. Push it one way and cheating slips through. Push it the other and honest candidates get punished for their kitchen lighting.
Too permissiveFair & defensibleToo aggressive
Failure mode A
The cheat gets through
Off-camera help, a second screen, a phone just outside the frame. The score is real to everyone reading it — and it is fiction.
Cost: a bad hire, and every honest candidate ranked below them.
Failure mode B
The honest candidate is flagged
A reflection mistaken for another person. A photo on the wall misread as a face. A sibling passing the door.
Cost: a qualified person rejected, and a decision you cannot explain.
Most systems optimise one and quietly pay for it with the other. A flag is not a verdict — and the only way to hold both ends of the dial is to make the system explain itself before it decides.
The eyes: CNNs for visual detection
Convolutional Neural Networks (CNNs) are the workhorses of image analysis. They’re designed to detect patterns - edges, shapes, and objects - and are used in everything from self-driving cars to facial recognition. For our proctoring system, we use YOLO (You Only Look Once), a state-of-the-art CNN architecture that identifies relevant elements in real time with high precision.
At Maki, we used two specialized CNNs: one to detect faces, and another to detect electronic devices such as phones, laptops, and monitors. The reason is simple - detecting a human face and detecting a screen involve very different visual cues. By training separate models for each, we gain accuracy and significantly reduce the risk of false positives.
Figure 02 · Inside the eyes
Two narrow models beat one that tries to see everything
Detecting a human face and detecting a screen involve very different visual cues. Training one model to do both blurs the line between them; training two keeps each decision boundary clean.
Model A — Faces
One question only
Trained exclusively on human faces. It answers how many, and where — nothing else.
Model B — Devices
One question only
Trained exclusively on electronics. It answers is there a device in this frame — nothing else.
Why it is built this way
Specialisation is a false-positive strategy, not a performance one. Each model holds a single, tightly drawn line, so when it fires, the signal means one specific thing — and the reasoning layer downstream knows exactly what it is being asked to interpret.
The brain: LLMs for contextual reasoning
While CNNs are excellent at identifying what’s in an image, they can’t tell you what’s really happening. They might see “two faces” or a “phone,” but they can’t determine intent. Is that second face a real person or a reflection? Is the phone being used or just sitting idle?
To bridge that gap, we introduced Large Language Models as a reasoning layer on top of the visual detectors. Our LLMs don’t just process object data—they analyze the entire image holistically, reasoning like a human reviewer would.
When a CNN flags something suspicious, the LLM takes over. It examines the image, applies clear rules, and explains its reasoning step by step. This process ensures that each alert comes with an interpretable, traceable justification—something traditional proctoring systems often lack.
Figure 03 · Why hybrid
Neither half is trustworthy on its own
Vision models and language models fail in opposite directions. That is precisely why they are used together — each one covers the other's blind spot.
| CNN vision modelsThe eyes · YOLO architecture | Large language modelsThe brain · contextual reasoning | |
|---|---|---|
| What it does | Scans each image and reports what is physically present — faces, phones, laptops, monitors. | Analyses the entire image holistically and reasons about what is actually happening. |
| Strength | Real time, high precision, tireless. Never gets bored on frame 40,000. | Judgement. Reasons the way a human reviewer would, and writes down why. |
| Blind spot | Cannot determine intent. A reflection and an accomplice are the same detection. | Has no eyes of its own — it needs something to point it at the moment that matters. |
| Alone, it produces | False positives — honest candidates flagged by their surroundings. | Nothing. There is no signal to reason about. |
| Its job here | Notice. Raise a signal — and stop there. | Explain. Turn the signal into a decision with a traceable justification. |
Detection without reasoning is noise. Reasoning without detection is blind. The hybrid exists because fairness lives in the gap between the two.
We use two LLMs in parallel: one focused on determining if multiple engaged people are in the frame, and another on assessing whether a detected device is actively being used. If no anomalies are detected, the system skips the LLM stage entirely, conserving resources and keeping the process smooth for candidates.
How the system makes decisions
The hybrid method operates in three stages:
- Detection: The CNN scans each image to identify faces and objects that could indicate cheating.
- Reasoning: If something unusual appears, the LLM evaluates it, applying structured logic to decide whether it’s a real concern.
- Decision: The system outputs an alert category and a concise justification.
Figure 04 · The architecture
Three stages, and only the last one decides
The eyes find things. The brain works out what they mean. Nothing becomes an outcome until both have spoken — which is what makes the result explainable rather than merely automatic.
Stage 01
Detection
The eyes — two YOLO models
One trained only on faces, the other only on electronic devices. Different visual cues, so they are kept as separate models.
ProducesRaw visual signals. No judgement, no conclusion.
Stage 02
Reasoning
The brain — two LLMs in parallel
One asks whether multiple engaged people are in frame. The other asks whether a detected device is actively being used. Each applies clear rules and explains itself step by step.
ProducesAn argument you can read, disagree with, and audit.
Stage 03
Decision
The outcome — with its reasons attached
Only now does the system act. It outputs an alert category and a concise justification for it.
ProducesA defensible call, not an anonymous red flag.
The quiet pathWhen stage one finds nothing unusual, the reasoning stage never runs at all — conserving resources and keeping the test smooth for the candidate. The overwhelming majority of sessions take this route.
The order matters more than the components. A signal that skips stage two becomes an accusation — and an accusation nobody can explain is the one thing a hiring team cannot defend.
This layered approach allows Maki’s system to move from raw visual signals to reliable, explainable decisions - a major leap forward from traditional proctoring models.
A closer look: when reasoning makes the difference
Consider a case where the camera detects two faces. In most systems, that would immediately trigger an alert. But our hybrid method looks deeper. The LLM checks if the second face is real, using cues like shadows, depth, and texture. It then determines whether that person is actually engaged with the candidate or just passing by. Only if both conditions are met does the system raise an alert.
Figure 05 · Worked example
One signal. Two candidates. Two different right answers.
This is the whole argument for putting reasoning between detection and decision. The vision model reports the same thing in both cases. The correct outcome is not the same.
Signal from the vision model
Two faces detected in frame — in most systems, this alone triggers an alert
Two conditions. Both must be met before an alert exists.
01Is the second face real? The model reads cues like shadows, depth and texture to separate a person from a reflection, a photo on the wall, or a screen behind them.
02Is that person engaged with the candidate? Someone genuinely participating behaves differently from someone just passing by.
No alert raised
A real face, yes — but a housemate crossing the doorway, not engaged with the candidate. The second condition fails, so nothing happens.
Under a threshold-only system this candidate is flagged — and has to argue their way out of it.
Alert raised
A real person, and one actively engaged with the candidate. Both conditions met — this is the case the system exists to catch.
And it arrives with its justification written out, not as a score a reviewer has to trust blindly.
The difference between the two columns is not a better camera or a tighter threshold. It is a system that asks what it is looking at before it decides what to do about it.

This reasoning dramatically reduces false positives. It ensures that candidates aren’t unfairly flagged for harmless situations, while real cheating is still caught with high confidence.
More than technology: building trust in assessment
Proctoring isn’t just a technical problem - it’s a matter of trust. Candidates deserve to know that they’re being evaluated fairly, and companies need confidence that their assessments reflect true ability.
By combining the precision of CNNs with the contextual intelligence of LLMs, Maki delivers a proctoring solution that safeguards both. It upholds test integrity without compromising the candidate experience, demonstrating that fairness and effectiveness can coexist.
In short, our hybrid approach turns monitoring into intelligent oversight - a system that sees clearly, reasons fairly, and earns trust at scale.
The principle
Integrity and fairness are not a trade-off. They are the same requirement.
A system that catches every cheat by suspecting everyone has not protected the assessment — it has destroyed the thing the assessment was for. Trust is what makes evaluation at scale possible, and it is built the same way in software as it is in a room: by being able to say exactly why you reached the conclusion you reached.
See what Maki Agents can do for you



