AI Can Mark Text. But Can It Mark Video?

Dr Rajeshwari Iyer
-
September 29, 2026

Assessment has always leaned heavily on the written word. Essays, short answers, multiple-choice papers - these are the formats AI marking tools have learned to handle well, and the progress in that space over the last few years has been genuinely impressive. But a huge amount of vocational and technical learning simply doesn't live on a page.

‍

Think about how you'd actually judge whether an apprentice has fitted their PPE correctly before starting work on a factory floor, or whether a hospitality student has the calm, confident manner needed to greet a guest at a hotel reception desk. Those are physical, behavioural, human judgements — the kind that traditionally require a human assessor watching a video, over and over, applying their expertise and experience to reach a fair, consistent decision.

This article rounds up the highlights of a recent webinar bringing together the people who designed, funded, and delivered that project. It will cover what worked, what proved genuinely difficult, and what it might mean for the wider assessment sector as multimodal, video-based evidence becomes a bigger part of how we measure competence. If you work in assessment, quality assurance, or qualification design and have ever wondered whether AI could ever get near judging something as human as a soft skill or a physical technique, this is worth fifteen minutes of your reading time.

The webinar, hosted by Tim Burnett of the Test Community Network fame, brought together four organisations at the heart of the project: Ufi VocTech Trust, the funder and long-standing champion of vocational technology innovation; SIAS, which develops engineering and mechanical qualifications and needed a way to assess practical skills and PPE compliance from video; CTH Awards (the Confederation of Tourism and Hospitality), that faced the challenge of judging softer skills, confidence, and interpersonal ability in hospitality candidates; and sAInaptic, the AI marking technology provider whose model was extended and tested across both use cases.

What makes this project interesting is precisely how different those two use cases are. SIAS needed the AI to make relatively objective judgements, grounded within the context of the apprentice's workplace: was the PPE put on in the right order, was a piece of equipment handled safely, did the candidate follow the correct safety protocol? That's the kind of task that, in principle, plays to AI's strengths - clear criteria, observable objects and actions, consistent standards to apply frame by frame.

CTH's challenge sat at the opposite end of the spectrum. Judging whether someone has the skills to work at a hotel reception desk isn't really an objective problem. It's about tone of voice, body language, warmth and patience with customers, the subtle read of a difficult guest interaction. These are the sort of nuanced, contextual judgement skills experienced human assessors develop over years and often struggle to fully articulate, let alone codify. Extending an AI marking model to have any credible opinion on that is a much harder technical and pedagogical problem.

Guest speaker Dr Rajeshwari Iyer, Co-Founder and CEO of sAInaptic, brought the technical perspective, explaining how the company's existing AI marking model, built and proven on text, was adapted to handle multimodal, video-based evidence. Deborah Hoggett of SIAS spoke to the practical realities of applying this to STEM and apprenticeship assessment, while Angela Hagenow of CTH offered the hospitality side, drawing on decades of experience in both hotel operations and academic assessment design. Jane Holmes of Ufi VocTech Trust rounded out the panel with the funder's view on why this project was worth backing, and what Ufi hopes the sector takes from it.

‍

Watch the full recording here, along with results on model performance and detailed profiles of all four panellists for anyone wanting to dig deeper into their backgrounds before tuning in.

‍

‍