Version 4 is live with Human Dialer and AI Workbench
AutoNurtureAutoNurture
Back to blog

AI in Utility Customer Service: The 90-Day Rollout Plan

Satisfaction with energy customer service is at an all time high. Among the customers who actually phoned their supplier in the last three months it fell five points. Here is the 90 day plan for putting AI on your busiest queues without touching the hard calls.

9 minPlaybook
A contact centre agent in a headset taking notes on a pad mid call, with a colleague on another call at the desk behind him.

Satisfaction with energy customer service is sitting at an all time high, 77 percent. Among the customers who actually phoned their supplier in the last three months, it fell to 81 percent from 86 percent in the previous wave.

Read that again. The people who never called are happier than ever. The people who called are less happy than they were.

Both numbers come from Ofgem most recent energy consumer satisfaction survey.

AI in utility customer service works best as a phased rollout: two weeks pulling a baseline, two weeks picking queues and signing off scripts, a 30 day pilot on two or three repetitive call types, then 30 days scaling what held. Humans keep vulnerability, complaints and negotiation. That is the whole plan.

This is written for contact centre and credit managers at energy, water, gas and telecom operators, and for the BPOs and collections agencies that run those queues for them. If you own an ACD dashboard and a headcount budget, keep reading.

What does AI in utility customer service actually handle?

Start with the calls that repeat. Not the sensitive ones.

The call types that survive an automation review in most utility contact centres look like this.

  • Meter reading submission and validation.
  • Balance queries, payment date changes and direct debit amount questions.
  • Bill explanation after a tariff or instalment change, the classic bill shock call.
  • Outage and restoration status updates during a spike.
  • Move in and move out, change of tenancy and final bill queries.
  • Day 5 and Day 15 payment reminders, plus promise to pay follow ups.
  • After hours and weekend overflow, when the alternative is voicemail.

The pattern is fixed path, high volume, low judgement. If resolving the call needs a policy exception, a vulnerability assessment or a negotiation, it belongs with a person.

A voice AI worker can answer inbound in seconds with no IVR menu, identify the account, resolve the fixed path cases and route everything else to an agent with the transcript and intent already on screen. AutoNurture.AI runs that on one shared dialer for AI and human calls, so the handoff lands on the same number and the same call log.

Days 1 to 14: pull the baseline before you touch anything.

Most failed pilots fail here. Nobody wrote down what normal looked like.

Pull six things and put them in one sheet.

  1. Abandonment rate by half hour, for your two busiest weekdays.
  2. Average handle time and first call resolution by call reason, top ten reasons only.
  3. Call volume by reason, ranked, with the share of total each one takes.
  4. Calls offered outside opening hours, and how many of those callers ever came back.
  5. Cost per contact, fully loaded, including supervision and QA.
  6. Promise to pay rate and kept promise rate, if you run collections in house.

Stop here for a second. Open your ACD and pull one number: abandonment between 09:00 and 11:00 on your busiest weekday.

If it is above 8 percent you have a coverage problem. Training will not fix it. More seats will help, and so will taking the repetitive half of that queue off your agents.

You cannot pilot your way out of a baseline you never measured.

Worth knowing what somebody else already publishes about you. In Great Britain, contact waiting time is one of four scored categories in the Citizens Advice supplier customer service comparison, and in the most recent published quarter the spread across suppliers ran from 1.8 out of 5 to 4.6 out of 5. Someone outside your building is already scoring your queue.

Days 15 to 30: pick two queues, not ten.

Two queues. Three at the absolute most.

Use five filters.

  • Volume above roughly 300 calls a week, so you get a readable signal inside 30 days.
  • A fixed resolution path that fits on one page.
  • No vulnerability screening needed to resolve it.
  • An outcome you can verify in a system of record: a meter read logged, a payment link clicked, a plan created.

An illustrative example, not a customer: a 40 seat water utility contact centre in northern Portugal. August, holiday roster, 22 percent abandonment by 10:40 on a Monday. Their top two call reasons are meter reading submission and instalment queries, together 41 percent of volume. Neither needs judgement. That is where the pilot goes.

Pick the boring queue first. The boring queue is where the volume lives and where a mistake costs you least.

Now the consent and GDPR work, as a checklist rather than a crisis.

  • Lawful basis for each outbound campaign, written down per market.
  • Call recording notice and retention window agreed with your DPO.
  • Disclosure at the top of the call that the caller is speaking to an AI, and an immediate route to a human on request. The European Commission framing of AI Act disclosure expectations is that people should be made aware when they are interacting with a machine.
  • Data residency and subject access request handling. AutoNurture.AI is EU hosted, with a DPA on request and configurable retention, which is the box most European utility procurement teams check first.

None of this is legal advice. Get your compliance team to sign the script and the retention policy before a single call dials.

Days 31 to 60: run the pilot like an experiment.

Thirty days. One queue at a time. A control group you did not touch.

  1. Route a fixed share of the queue, 20 to 30 percent, to the AI worker. Leave the rest with your agents as a control.
  2. Set escalation triggers before go live: frustration, a request for a human, a vulnerability cue, any balance above your threshold, any compliance keyword.
  3. Review 20 transcripts a day. Not a sample report. Actual transcripts, read by the supervisor who owns the queue.
  4. A/B the opener. The first eight seconds decide whether the call resolves or transfers.
  5. Hold a 30 minute weekly review with one scorecard and one owner.

The scorecard is seven lines, the same every week.

  • Containment rate: calls fully resolved with no transfer.
  • Transfer rate, split into requested transfers and triggered transfers.
  • First call resolution on AI handled calls, measured the way you measure it for agents.
  • Abandonment on the human queue, which should fall as the AI absorbs volume.
  • A one question post call score on AI handled calls.
  • Promise to pay rate and kept promise rate, if the queue is collections.
  • Cost per contact against your baseline number.

Live transcript analysis is what makes a weekly review possible at all. Sentiment, intent and risk signals surfacing during the call, across AI and human conversations, means your supervisor reads the exceptions instead of the whole log. That is what the AI Workbench and live call analysis are for.

Your go or no go gates at day 60: containment above 50 percent on the chosen queue, triggered transfer rate under 15 percent, post call score within two points of your agent baseline, zero compliance incidents. Miss a gate, extend two weeks and change one thing. Miss it twice, drop that queue and take the next one on the shortlist.

Days 61 to 90: scale what held, keep humans on the hard calls.

Now widen it, in this order.

  1. Take the pilot queue from 30 percent to full volume.
  2. Add after hours and weekends on the same queue, because that is coverage you were already sending to voicemail.
  3. Add the second queue from your day 15 shortlist.
  4. Add proactive outbound on the same call types: payment reminders on cadence, appointment confirmations, instalment change notifications.
  5. Write the escalation playbook into your QA framework so new agents inherit it.

Volume to the AI. Judgement to your agents. If your rollout cuts headcount, you built the wrong plan.

The point is that the person who is good at a difficult arrears conversation spends the day having difficult arrears conversations instead of reading meter numbers back to people.

Seasonal spikes are the clearest case. A cold snap, a tariff change, a billing run that goes wrong, and volume triples for nine days. You cannot hire for nine days. You can put the repetitive half of that spike on a worker that answers every call on the first ring and keep your team on the calls that need a person. AutoNurture.AI covers this sector in its utilities and energy playbook.

Where this rollout usually goes wrong.

Four failure modes, all avoidable.

  • Starting with your hardest queue. Complaints and vulnerable customer calls are the worst possible pilot and the most tempting, because that is where the pain is loudest.
  • No control group. Without one, every improvement gets argued about for a quarter.
  • No frustration trigger. A caller who wants a human and cannot reach one costs you more than the call was worth.
  • No named owner. A pilot owned by a committee dies in week three.

What to do next.

Pick your queue. Pull the six baseline numbers. Write the one page script.

Then hear what these calls actually sound like before you commit anything. AutoNurture.AI publishes recorded utility demo calls with transcripts in German, Portuguese, Italian and English.

When you want to see it against your own queue and your own stack, book a demo. Twenty minutes, in your language, with a worker on a live call.

Frequently asked questions

How long does it take to roll out AI in a utility contact centre?

Ninety days is realistic for one or two queues: two weeks of baseline, two weeks of queue selection and compliance sign off, a 30 day pilot on a fixed share of one queue, then 30 days of scaling. Launching six queues at once is what stretches a rollout to nine months.

Which utility calls should stay with human agents?

Anything that needs judgement or care: vulnerability cases, complaints and escalations, disputed bills, affordability conversations, anything above your balance threshold, and any call where the customer asks for a person. Route those with the transcript and intent attached so the agent does not restart the conversation.

What containment rate should we expect from a first pilot?

Set a gate rather than a forecast. On a fixed path queue such as meter reading submission or balance queries, containment above 50 percent with a triggered transfer rate under 15 percent is a reasonable day 60 gate. Arrears and complaints queues sit far lower, which is exactly why they are not pilot queues.

Do we have to tell customers they are speaking to an AI?

Disclose it, every time, and give an instant route to a human. Confirm the wording with your compliance team for each market.

How does this work across Portuguese, German, Italian and English queues?

Multilingual voice is standard now. AutoNurture.AI answers in 12 or more languages with native voice rather than translated scripts, and logs the call in your working language. Run the pilot in one language first, then add the next once the scorecard holds.

Will an AI voice worker reduce our headcount?

That is the wrong target. An AI absorbs the repetitive half of the queue: meter reads, balance checks, reminder calls, after hours overflow. What changes is what your agents spend the day on, and how many callers hang up before anyone picks up.

What should we measure every week during the pilot?

Seven lines: containment, transfer rate split by reason, first call resolution, abandonment on the human queue, post call score, promise to pay and kept promise rate if it is a collections queue, and cost per contact against baseline. Same seven every week, one owner, thirty minutes.