Blogs, News & Insights

Why your current live training evaluation doesn't scale

Written by Kayleigh Tanner | 22 September 2026, 08:40:45 Z

Human observation can tell you a lot about an individual training session. But when your organisation delivers hundreds or thousands of live training sessions a year, it can only ever tell you a small part of the story.

Despite the wealth of tech-first approaches to workplace learning, organisations continue to invest heavily in live training. Instructor-led training (ILT), virtual instructor-led training (VILT), workshops and practical sessions remain an important part of workplace learning, with 98% of organisations still including some form of live training in their learning strategies. Those organisations have teams of trainers responsible for turning learning content and instructional frameworks into memorable, engaging experiences that help people understand, practise and apply new skills.

That makes training quality important. But measuring it consistently hasn’t always been... well, consistent.

Most organisations already have some form of training evaluation in place. Learners complete feedback surveys, managers occasionally observe trainers, sessions may be recorded for later review, and in large organisations, there may even be a quality team to evaluate training delivery against established criteria.

These approaches all have a place in live training evaluation. The problem is what happens when you try to scale them.

 

The limits of human-led training evaluation

A manager can sit in on a training session and assess how effectively a trainer communicates, engages learners and responds to questions.

They might notice that a trainer explains complex concepts particularly well. They might identify that another trainer struggles to keep discussions focused. They might see a third trainer adapt their approach when learners are confused.

This kind of observation can provide valuable insight, but it’s also time-consuming. If a manager wanted to observe every single training session, they wouldn’t have time to do anything else.

Realistically, managers can’t observe every session delivered by every trainer, and even if sessions are recorded, someone still needs to find the time to watch and evaluate them. As the number of trainers, sessions and locations increases, the proportion of training that can be evaluated in this way quickly becomes very small. A manager might be able to observe one training session a month, but what’s happening in the other 19 sessions they’re not able to evaluate?

The result is a familiar problem for L&D teams: the more training you deliver, the harder it becomes to see what is actually happening across that training.

 

A simple example

Imagine a large multinational organisation with 100 trainers, each delivering two live sessions a week.

That's around 10,000 training sessions a year.

If a manager observes each trainer once a quarter, they're directly assessing just 400 sessions.

That's only 4% of the training being delivered.

And that assumes every observation happens as planned, every session is representative of the trainer's usual performance and there’s enough management capacity to provide meaningful feedback afterwards.

This is a relatively generous example – the reality is likely to be even more limited. Managers have limited capacity, evaluating the quality of live training delivery isn’t a top priority and even when training sessions are observed, very little usually happens with that evaluation data.

 

From small samples to big assumptions

The problem with limited observation isn't necessarily that the observations themselves are inaccurate (though unconscious bias or a whole host of other factors can mean that trainers aren’t assessed consistently or entirely objectively).

The real problem is what organisations are sometimes asked to conclude from the very limited data available to them.

A trainer might deliver an excellent session on the day they're observed. Does that tell you how they typically perform, or are they on best behaviour because they know their work is being judged?

Another trainer might have a difficult session because the technology fails, the group is unusually challenging or the subject matter requires more explanation than expected. Does that single observation accurately represent their overall capability?

Without a broader sample, it’s extremely difficult to extrapolate trends and patterns from these infrequent, individual snapshots.

This matters because effective training delivery shouldn’t just look at whether a trainer covered the right content or whether the learners enjoyed the session. Instead, live training evaluation needs to consider how the trainer delivered the session, which means being able to answer questions like:

  • Did they explain concepts clearly?
  • Did they notice when people were struggling?
  • Did they adapt their approach when necessary?
  • Did they create opportunities for learners to participate?
  • Did they provide useful feedback during practical activities?

These are behaviours that can have a significant impact on the learning experience, but they're difficult to assess consistently when evaluation depends on someone a) being there to observe them, and b) being able to spot those behaviours consistently across different delivery styles and session types.

 

Live training evaluation has a scalability problem

Traditional training evaluation often combines several sources of information:

Learner feedback

Learner surveys can reveal how people felt about a session, whether they found it useful and how they perceived the trainer.

But learner feedback doesn't necessarily explain why a session worked or didn't work. A popular trainer isn't automatically an effective trainer, and a challenging subject can result in lower learner satisfaction, even when the trainer objectively delivered it well.

 

Manager observation

Direct observation allows an experienced L&D professional or manager to assess the training in context.

But observation takes time. The number of training sessions that can be assessed is constrained by manager availability across both the session observation itself (whether that’s attending the live session or watching a recording) and the capacity to document their findings and communicate feedback to the trainer.

 

Completion and performance data

Completion rates, knowledge checks and other outcome measures can show whether learners finished training or demonstrated knowledge afterwards.

These activity measures (sometimes called vanity metrics) useful in their own way, but they don't necessarily reveal what happened during the training that contributed to those outcomes. They tell us that the training session got learners from point A to point B, but not how they got there, which makes it difficult to replicate good sessions or change the approach to avoid future ineffective training.

None of these measures is inherently problematic, but the challenge comes from trying to use them to understand the quality of an entire training programme when each only provides a limited view.

 

The training evaluation visibility gap

This creates a gap between what organisations want to know and what traditional training evaluation can realistically tell them.

 

What L&D wants to know

What traditional evaluation can provide

How consistently are trainers delivering effective sessions?

Feedback from a limited sample

Which trainers have particular strengths?

Observations and learner perceptions

Which skills need development across the team?

Findings from individual assessments

Are trainers improving over time?

Periodic comparison

What behaviours are contributing to effective training?

Subjective observations and outcome data

 

The information exists, but it is often fragmented across surveys, observation notes and recording analyses.

And because human observation is difficult to scale, the evidence available to L&D teams can be much narrower than the training activity it is supposed to represent.

 

From checking single sessions to understanding patterns

This doesn't mean human observation should disappear.

Experienced L&D professionals bring context and judgement that technology cannot replace. They can understand the nuances of a training environment, interpret what happened and decide what action should follow.

The opportunity is to give them more evidence to work with.

Instead of relying on occasional observations to understand trainer capability, imagine being able to look across a much broader sample of real training sessions.

You could start to identify patterns:

A trainer who consistently creates strong learner engagement.

A trainer who explains technical processes clearly but rarely checks understanding.

A team that is strong at facilitating discussion but could improve how it adapts delivery to learner needs.

A trainer whose skills are developing following targeted learning and coaching.

This moves training evaluation beyond the question of “How did this session go?”

Towards:

“What can we learn from the training we're delivering at scale?”

 

The next step: measuring the skills behind training quality

The challenge isn't that organisations don't know what effective training looks like.

Most L&D teams already have frameworks, competency standards and experienced professionals who can define good training delivery.

The challenge is measuring whether those skills and behaviours are actually being demonstrated consistently across a training programme.

That's where AI introduces a new possibility.

Rather than relying solely on someone watching a session, AI can analyse the conversations taking place during training and identify evidence of specific training skills and behaviours.

It opens up the possibility of making trainer capability more visible across a much larger proportion of training activity, while still leaving human judgement at the centre of the evaluation process.

In our next post, we'll explore how AI-powered skills intelligence can be used to measure the behaviours behind effective training delivery, and what this could mean for the future of training quality assurance. Watch this space...