Science
AI-powered skill assessment: the science behind the new paradigm
A New Paradigm for Skill Assessment
March 12, 2025
Ben Chino

AI-powered skill assessment is the use of artificial intelligence to score, and increasingly to generate, the tests employers use to measure what a candidate can do. As skill assessment moves into a new era, the convergence of artificial intelligence and psychometrics is reshaping how we evaluate human abilities. At Maki People, we combine AI technologies with a range of psychometric techniques to ensure reliability, fairness and validity in AI-powered assessments.
This article sets out our current methodology and the research still in progress. It explains how AI-scored tests, whether in multiple-choice, written or spoken formats, can match or eventually exceed human raters in accuracy, objectivity and fairness, and what has to be checked before anyone should take that on trust.
From multiple choice to AI-scored open text
Recruitment assessment has historically relied on structured formats such as multiple-choice questions, which offer limited insight into a candidate's critical thinking and job-relevant abilities. These formats constrain how candidates can demonstrate what they know, and often produce an incomplete picture of their potential.
Open-ended responses have traditionally been scored by human raters. That process is slow, and human scoring carries inherent limitations: subjectivity, inconsistency between recruiters, time pressure, and the difficulty of scaling to high-volume hiring. Psychometric frameworks such as Item Response Theory (IRT) and Structural Equation Modelling (SEM) ensure the robustness of assessments, but they have mainly been applied to structured test formats that do not fully capture a candidate's problem-solving approach or communication style.
By incorporating speech recognition and text-to-speech synthesis models alongside text-based AI evaluation, we extend beyond the standard assessment experience to a wider range of candidate responses, written and spoken. Speech-to-text (STT) models transcribe candidate speech accurately into text, while text-to-speech (TTS) technology makes automated interviewer interactions more dynamic and natural.
What generative AI changes in skill assessment
With the advent of generative AI (Gen-AI), we are witnessing a paradigm shift in recruitment assessment. The technology lets organisations move beyond traditional approaches in several ways:
- Gen-AI utilises large language models (LLMs) to evaluate complex, open-ended responses that demonstrate job-relevant competencies.
- Gen-AI leverages LLMs and natural language processing (NLP) to assess the power skills crucial for workplace success.
- Gen-AI applies consistent scoring criteria across thousands of applications without fatigue.
- Gen-AI analyses candidate reasoning in real time, giving hiring teams immediate insight.
- Gen-AI employs LLMs to generate content on the fly for bespoke, client-specific tests.
The questions AI scoring has to answer
This transition introduces challenges that any AI-powered skill assessment must address before it is used to make decisions about people:
- Are scores generated by LLMs comparable to the evaluations of experienced hiring managers?
- Does AI scoring introduce or amplify systematic biases across demographic groups?
- Can AI-generated content achieve psychometric equivalence to items written by professional item writers?
To answer these questions, Maki People applies advanced statistical methods to evaluate AI-generated scores rigorously, so that our recruitment assessments are not only efficient and scalable but scientifically sound and fair for candidates.
How we validate AI scoring against human raters
Inter-rater agreement and calibration
A crucial step in our validation process is assessing the agreement between AI-generated scores and human raters, the same starting point set out in the established framework for evaluating automated scoring used in educational measurement. We use inter-rater reliability (IRR) metrics such as intraclass correlation coefficients (ICC) and Cohen's kappa to evaluate the consistency between AI and expert scorers.
To investigate potential discrepancies between LLM and human scoring patterns in depth, we employ:
- Bland-Altman analysis to identify systematic differences or biases in scoring across the range of performance levels.
- Statistical analysis with demographic variables to detect whether AI-human score differences are associated with specific candidate characteristics.
- Regression analysis to determine whether factors beyond candidate ability influence scoring discrepancies.
- Error pattern analysis to identify the specific types of response where AI and human ratings tend to diverge.
This approach ensures that our AI scoring models are calibrated to human expert judgement while maintaining fairness and accuracy across diverse candidate populations. For more detail on how we evaluate LLMs, see Building Transparent AI: Why Trust Is Non-Negotiable.
Later in the validation process we run further analyses to check that the assessment system as a whole functions fairly across demographic groups. The initial AI-human comparison focuses on establishing scoring equivalence.
Evaluating speech-to-text models with expert review
One of the critical aspects of AI-driven skill assessment is the accurate transcription and interpretation of spoken responses. At Maki People, we evaluate speech-to-text models rigorously for accuracy, consistency and applicability across diverse candidate populations. Many off-the-shelf STT solutions claim high accuracy, but their real-world performance varies significantly across accents, speech patterns and background noise.
To validate STT models we use a dual-layered evaluation process:
- Expert human review. Transcriptions generated by STT models are compared against expert or native-speaker transcriptions, using metrics such as Word Error Rate (WER), Phoneme Error Rate (PER) and Semantic Error Rate (SER). This checks that the model captures key linguistic nuances.
- Empirical evaluation. We assess whether transcription errors systematically affect test scores, particularly for non-native speakers or for people with specific accents or speech impairments.
- Impact on automated scoring. STT outputs are the input to the LLMs that perform automated scoring. Where our analysis identifies languages or dialects in which STT performance is inadequate, we exclude those languages from our automated speech assessment pipeline to protect assessment integrity.
Pronunciation assessment across languages
Beyond transcription, AI-powered assessments must reliably evaluate pronunciation quality, particularly in multilingual settings. Many AI-driven pronunciation models claim to support multiple languages; their effectiveness across different linguistic contexts varies significantly.
To assess the performance of pronunciation assessment models, we conducted a structured evaluation across multiple languages:
- Native speaker review. Native speakers rated the accuracy of pronunciation scores as low accuracy, accurate or high accuracy.
- Independent expert review. A second layer of evaluation was carried out by external linguistic experts, for a broader perspective.
- Comparative analysis. We compared the ratings from native speakers and experts to identify systematic discrepancies and to assess the reliability of the model across languages.
These findings underline the importance of rigorous validation when integrating pronunciation models into AI-driven evaluation. Systematic biases, especially across diverse linguistic groups, must be monitored carefully to maintain fairness and accuracy in automated speech assessment. We apply the same dual-layered approach to our AI-based scoring models, so that both transcription quality and pronunciation evaluation are validated before deployment.
Psychometric evaluation of AI-powered assessments
Despite our use of AI throughout the recruitment and assessment process, we remain committed to psychometric rigour. We employ well-established psychometric frameworks so that our assessments meet the highest scientific standards. The full description of how Maki assessments are built, validated and audited is on our science page; the three frameworks that matter most are:
- IRT for item-level analysis. Item Response Theory focuses on the performance of individual items, ensuring that each question is psychometrically sound and contributes meaningfully to the assessment.
- SEM for construct-level validation. Structural Equation Modelling evaluates the overall structure and validity of the assessment, ensuring that the test as a whole measures the intended latent constructs accurately and consistently.
- DIF for fairness analysis. Differential Item Functioning analysis examines whether items perform equivalently across demographic groups, detecting and addressing potential sources of bias so that every candidate is evaluated equitably regardless of background.
Together, IRT and SEM ensure that our AI-powered skill assessments satisfy rigorous psychometric criteria for reliability, validity and fairness, the scientific integrity expected in high-stakes assessment. Combining cutting-edge AI with established psychometric methodology lets us keep assessment quality while gaining the efficiency and scalability of AI-driven evaluation.
Example: the Problem Solving (Advanced) test
One example of this process is the Problem Solving (Advanced) test, a widely used assessment of critical thinking, logical reasoning and analytical skills, all essential for evaluating complex information and making data-driven decisions. To ensure scientific rigour and fairness, we applied:
- Item Response Theory to assess item difficulty and discrimination, confirming that the test differentiates effectively between high- and low-performing candidates. The test demonstrated high internal consistency (0.87) and strong psychometric properties.
- Differential Item Functioning analysis to verify that items perform equally across demographic groups (gender, age, ethnicity). No significant bias was found.
This is one example of how we evaluate our assessments so that they are scientifically robust, fair and predictive of real-world performance, while keeping the efficiency of AI-driven evaluation.
LLMs for assessment content generation
Our focus so far has been AI-driven scoring, but large language models offer value in assessment content generation that extends well beyond evaluation. At Maki People, we use LLMs across the assessment creation process, producing content that is diverse and relevant before it goes through psychometric testing. This is how bespoke tests are built for Ken, our long-form structured assessment agent.
The benefits of LLM-powered content generation in skill assessment include:
- Rapid creation of diverse assessment items across domains, skill levels and formats.
- Difficulty calibration, through carefully crafted questions and problems that target different levels of difficulty.
- Domain-specific scenario generation that simulates real-world workplace challenges with authentic context.
- Culturally sensitive, globally relevant content that minimises geographic or cultural bias.
- Automatic distractor generation for multiple-choice assessments, with plausible yet clearly incorrect options.
Our research shows that LLM-generated assessment content, when properly validated with psychometric techniques, achieves comparable or superior item discrimination and reliability compared with traditionally developed assessments. The ability to iterate and refine content rapidly also allows continuous improvement of assessment quality.
Combining human expertise with LLM capability creates a powerful synergy: psychometricians can focus on validation and refinement rather than initial content creation, which dramatically accelerates the assessment development lifecycle while maintaining scientific rigour.
The future of AI-driven skill assessment
Our work is far from complete. As AI evolves, so do our methods for refining and validating automated scoring algorithms. Key areas of ongoing research include:
- Adapting Gen-AI models for multilingual skill assessments to ensure fairness across linguistic groups.
- Developing adaptive testing mechanisms, where the algorithm selects test items dynamically based on real-time performance.
- Expanding beyond textual analysis into conversational speech and multimodal assessments, where AI evaluates not just written responses but spoken communication in a conversational manner.
- Investigating Gen-AI's ability to provide real-time, personalised feedback to enhance learning outcomes.
Frequently asked questions
Is AI scoring as reliable as human raters? That is the first thing we test rather than assume. We measure agreement between AI and expert scorers with inter-rater reliability metrics such as ICC and Cohen's kappa, then use Bland-Altman, demographic, regression and error-pattern analyses to find where and why the two diverge, and calibrate the model to expert judgement.
How is an AI-powered skill assessment checked for bias? At two levels. Scoring equivalence is checked first, by testing whether AI-human score differences are associated with candidate characteristics. The assessment as a whole is then analysed with Differential Item Functioning to confirm that items perform equivalently across demographic groups.
Can test questions written by an LLM be trusted? Only once they have passed the same psychometric testing as human-written items. Our research shows that LLM-generated content validated this way achieves comparable or superior item discrimination and reliability.
A new paradigm for skill assessment
At Maki People, we are pioneering a new paradigm for skill assessment, one where AI scoring is not just automated but scientifically validated and psychometrically rigorous. Our combination of AI and psychometric modelling ensures that skill assessments are fair, accurate and scalable, helping employers, educators and candidates make data-driven decisions in an evolving workforce.
See what Maki Agents can do for you



