Scoring

How to score a CSR call so the score means something

Most call scorecards produce a number that correlates with nothing. The fix is not a longer rubric. It is scoring fewer things, choosing them on whether they plausibly cause a booking, and proving two people score the same call the same way.

A scorecard earns its place if a rep who scores higher books more. That is the entire test, and most scorecards fail it — not because the criteria are wrong in some moral sense, but because they were chosen for being easy to observe rather than for being connected to the outcome.

What most scorecards actually measure

Greeting used. Name captured. Tone friendly. Company name stated. Thanked the caller. These are real behaviours and they are trivially scoreable, which is why they end up on the form. They are also nearly uncorrelated with whether the job got booked, because almost every rep does almost all of them almost every time. A criterion that 96% of calls pass carries no information — it cannot separate anybody.

Three tests for a criterion worth scoring

Observable. Two people listening to the same recording can agree whether it happened, without knowing the outcome. "Built rapport" fails this. "Asked what the customer had already tried" passes.

Plausibly causal. There is a mechanism by which doing it more makes a booking more likely. Offering a specific appointment window rather than promising a callback has a mechanism. Saying the company name twice does not.

Controllable. The rep can do it on demand. Scoring people on whether the caller was in a hurry measures the queue, not the rep.

Criteria that survive all three tend to be few — often six to ten — and tend to be about the shape of the call rather than its manners: whether the problem was established before a price was discussed, whether an appointment was actually offered, whether an objection was answered or absorbed, whether the call ended with a committed next step or a vague one.

Score stages, not impressions

The most useful structure is not a list of adjectives but a sequence of gates the call either reached or did not: qualified, need established, appointment offered, objection handled, commitment obtained. Each is binary and observable.

The payoff is diagnostic rather than evaluative. When a rep's calls consistently die at the same gate, you know what to work on. A composite score out of a hundred tells you somebody is at 71 and gives you nothing to do about it.

Sample properly, or do not bother

Scoring the calls a manager happens to remember produces a picture of the calls a manager happens to remember. Two rules fix most of it.

Random within strata. Draw calls at random, but in proportion across lead source and job type, so a rep who fields more emergency work is not compared against someone fielding maintenance enquiries. Mix is the single biggest confound in per-rep comparison and sampling is where you can cheaply remove it.

Include the calls that did not book. Obvious when stated, routinely skipped. A scorecard built only from bookings cannot tell you what distinguishes a booking.

The step nearly everyone skips

Calibration. Have two people score the same fifteen calls independently, then compare criterion by criterion. Anywhere they disagree, the rubric is ambiguous — not the scorers. Rewrite that criterion until two people reach the same answer without conferring.

Until you have done this, you do not know whether your scores measure calls or measure scorers, and any per-rep ranking drawn from them may be a ranking of who was scored by whom. It is a couple of hours of work and it is the difference between a scorecard that supports a coaching conversation and one that starts an argument.

What this is for. Scoring is how you find out where in the call the booking is lost. It does not tell you what that loss is worth — that needs the conversion gap priced against your own job values, which is what the method does and what a diagnostic delivers.