Science

Building transparent AI: why trust in hiring is non-negotiable

Why Trust Is Non-Negotiable

March 12, 2025

12 mins

Maxime Legardez-Coquin

Building transparent AI: why trust in hiring is non-negotiable

Transparent AI in hiring means an assessment system whose scores can be explained, checked against human judgement, tested for bias and audited by someone outside the company that built it. At Maki People, we believe AI should empower, not obscure. When AI-driven assessments shape careers and businesses, trust is not just an ethical stance. It is a business imperative. Here is what transparent AI means in hiring, and what we are actually doing to achieve it.

What is transparent AI in hiring?

An AI hiring tool is transparent when a recruiter, a candidate or a regulator can answer four questions about it. What does it measure? How accurate is it compared with a qualified human? Does it treat demographic groups the same? Who checks it, and how often? An opaque system, often called a black box, produces a score and nothing else. A transparent one produces a score, the reasoning behind it, and the evidence that the reasoning holds.

Explainable AI is the part of this that concerns the individual decision: why did this candidate get this score? AI transparency is the wider commitment: publishing how the models were chosen, how they were validated, how fairness was tested and how humans stay involved.

Why AI transparency matters to employers and candidates

Transparent AI is not just an ethical choice. It is a competitive advantage. Businesses need to trust the tools they use to assess talent, and candidates deserve clarity on how decisions are made. At Maki People, we are committed to developing AI that provides insight into its decision-making, so hiring teams can rely on data-driven evaluations while candidates receive accurate, objective and fair scores that reflect their true potential.

Our progress so far:

  • Evaluating AI predictions against expert human ratings to ensure alignment and accuracy
  • Conducting internal audits to refine decision-making transparency
  • Continuously evaluating and selecting the best foundation models for reliable, accurate scoring
  • Working towards explainable AI models that give clearer reasoning behind candidate evaluations

How we ensure scoring accuracy

Precision in candidate evaluation drives everything we do. We have moved from single-choice questions to more insightful open-ended formats, and researched extensively which AI technologies understand candidate responses most accurately.

Testing semantic models on open-ended answers

Our research team rigorously tested multiple semantic analysis approaches, including BERT-based models, MiniLM-L6, MPNet, BGE embeddings and OpenAI's GPT-4 models, to find the best technology for scoring open-ended responses. The testing revealed clear performance differences, which guide our implementation decisions.

For example, a sentence embedding approach using models such as Multilingual MPNet showed significant improvements over traditional word-based methods, reaching up to 67% accuracy with a natural distribution of similarity scores that discriminates better between response qualities. Word embeddings produce artificially high similarities across almost all responses; sentence embeddings show a more realistic evaluation pattern, allowing meaningful differentiation. For greater accuracy still, our testing of OpenAI's GPT models delivered 81 to 82% accuracy and 95% precision in identifying incorrect answers, a substantial improvement over baseline approaches. The implementation includes robust error handling, automated retries and progress tracking, so results stay reliable whichever model is used.

The technical advantages are clear: significantly higher precision than traditional methods, more accurate identification of genuinely similar responses, and results that align better with expert human judgement. The set-up also lets us balance accuracy, speed and cost against the needs of a specific assessment. We keep refining the models through ongoing research and validation, which means more meaningful insights and fairer evaluations for every candidate, whatever their background or response style. You can read more about how our assessments are built and validated on the Maki science page.

Language proficiency at CEFR level

Beyond the model evaluations we run for situational judgement tests, we also test language proficiency assessment across LLMs with varying parameters to make sure the evaluation is robust. Our comparative analysis of seven OpenAI models, Claude models and DeepSeek across multiple prompts achieved up to 90% exact match in CEFR level predictions (A1 to C2). Testing across different use cases in this way keeps accuracy consistent whatever skill is being evaluated.

Fair and bias-tested AI: what our validation study found

AI-driven hiring should level the playing field, not reinforce bias. Accuracy is crucial, but fairness is equally important. Our research evaluated the reliability, validity and fairness of LLM assessments and compared AI scores with human raters across soft skills including result orientation, team orientation and adaptability. The findings showed strong alignment between AI and human evaluation, and Bland-Altman analyses confirmed minimal systematic bias. Statistical differences between specific demographic groups were small, typically less than one point, rather than practically meaningful.

{{BANNER}}

Agreement with human experts

Our recent validation study examined how our AI evaluates three critical soft skills: result orientation, team orientation and adaptability. It compared AI evaluations against ratings from multiple qualified industrial-organisational psychologists using a structured scoring grid. Intraclass correlation coefficients showed good to excellent agreement against standard reliability benchmarks.

Agreement between AI and the human expert consensus:

  • Result orientation: 91%
  • Team orientation: 92%
  • Adaptability: 89%

Correlation between AI evaluations and human expert ratings:

  • Result orientation: r = 0.86
  • Team orientation: r = 0.78
  • Adaptability: r = 0.73

Demographic fairness testing

We rigorously evaluated the LLMs for potential bias across demographic groups:

  • Gender: statistical testing found no significant scoring differences between male and female candidates
  • Age: no significant differences in scores between younger and older candidates
  • Systematic bias: Bland-Altman analyses indicated minimal systematic bias in AI evaluations

How we keep improving

We remain committed to fairness and transparency by:

  • Expanding our training data to represent diverse candidate pools
  • Iteratively improving our scoring methods
  • Including more demographic groups in fairness testing
  • Validating assessments on established statistical principles and psychometric techniques

This transparent approach means candidates are evaluated on their true abilities, not irrelevant factors, while organisations benefit from the efficiency and scale of AI-assisted evaluation.

Compliance and ethics: staying ahead of AI hiring regulation

Regulatory frameworks around AI in hiring are evolving fast. Companies that fail to adopt transparent AI risk compliance problems and legal consequences. At Maki People, we are proactively aligning with global AI ethics standards and working with independent auditing companies to keep improving our compliance practices.

Our commitment to fairness has been assessed through an independent audit by Holistic AI, a respected AI governance platform. Their assessment of several of our tests delivered encouraging results: all demographic groups evaluated were treated fairly in the hiring process. Applying the standard four-fifths rule, a key metric for identifying potential hiring discrimination, they found no concerning differences in selection rates between demographic categories. This held whether examining individual characteristics or intersectional combinations, confirming that our assessments meet rigorous fairness standards. Our current audit results and certifications are published on the Maki security and bias audits page.

Our ongoing efforts:

  • Subjecting our AI assessment models to independent audits and continuing to enhance our fairness metrics, so the technology stays unbiased and compliant with evolving regulation as the platform scales
  • ISO 27001 certification, demonstrating our commitment to information security and regulatory compliance
  • Aligning with the EU AI Act, GDPR, EEOC guidance and emerging AI hiring regulatory frameworks
  • Enhancing internal AI documentation to facilitate compliance audits
  • Regularly reviewing AI ethics best practice to ensure responsible development
  • Strengthening candidate privacy and data protection protocols

Human-in-the-loop validation

AI offers scale and consistency, but human expertise remains essential for maintaining assessment quality. Our research shows that combining AI capabilities with human oversight creates a more robust evaluation system than either approach alone.

Our technical approach:

  • Structured validation protocol: a systematic human review process in which experts evaluate a significant sample (around 10%) of AI-scored assessments going forward, to confirm ongoing accuracy and identify any scoring patterns that need calibration
  • Edge case detection: our monitoring system automatically flags unusual scoring patterns and edge cases for expert review, improving accuracy by surfacing the cases where AI confidence is lowest
  • Continuous learning framework: a feedback loop in which human reviewer corrections are documented
  • Audit documentation infrastructure: a proprietary tracking system that keeps comprehensive records of every assessment decision, enabling detailed fairness analyses across demographic factors and full transparency for compliance requirements

This human-in-the-loop approach complements our AI: automation for efficiency, human judgement as the foundation of our methodology.

Frequently asked questions

What is explainable AI in hiring? Explainable AI in hiring means a system that can give a human-readable reason for each score: which answers, criteria or skills drove the result. It is one component of transparent AI, alongside validation against human raters, fairness testing and independent audit.

How is an AI hiring tool tested for bias? By comparing outcomes across demographic groups. Common methods include the four-fifths rule on selection rates, statistical tests for score differences by gender and age, and Bland-Altman analysis for systematic bias between AI and human scores.

What does the EU AI Act mean for AI in recruitment? AI systems used in recruitment and selection fall into the high-risk category of the EU AI Act, which brings obligations around transparency, human oversight and documentation. Our EU AI Act guide for talent acquisition leaders sets out what to do before the rules apply.

Final thoughts: the future of AI in hiring starts with trust

At Maki People, transparency is not a feature. It is a journey. We keep improving our AI models to enhance explainability, fairness and compliance while maintaining human oversight. By prioritising responsible AI development, we help companies hire with confidence and integrity, and give candidates valuable insight into their potential for career growth. In a world where AI is reshaping recruitment, trust is the currency of success.

See what Maki Agents can do for you
Experience how Maki’s AI agents simplify, speed up, and elevate your hiring
Request a demo
See what Maki Agents can do for you

See what Maki Agents can do for you

Request a demo