The Agent Operations Method
Practices for running a support operation that outlives the people who built it.
Most support teams have good tools and struggle with the operations underneath them. The helpdesk is configured, the bot is live, the dashboard refreshes every morning, and the weekly report goes out on time. Still, nobody can say what good work looks like, who owns the queue, or which capability is failing when a number moves.
That distance between having tools and having an operation is where this method lives. It is written for the person accountable for the work rather than the person reading the report.
It assumes your team is part human and part agent, and that both need the same treatment. It assumes you are somewhere on a ladder that runs from reactive to optimized, and that you can only stand on one rung at a time. Skipping rungs is the most common way operational programs fail. Instrumentation laid over an undefined standard produces numbers that get argued about for a quarter and then quietly ignored.
None of this is finished. Operations change, evidence accumulates, and the method changes with it.
Principles
These are the beliefs underneath everything else. They are opinionated on purpose. An operation that holds no position produces no standard.
- 01
Competency is the unit
Every outcome you care about traces back to something a person or an agent can or cannot do. Tools get replaced and staff turns over. The map of what your operation is capable of is the thing worth maintaining.
- 02
Build the standard, then read against it
Most teams try to measure quality before writing down what quality is, which produces scores nobody trusts. If the standard does not exist, write it. Reading comes second.
- 03
Variance is the signal
An average tells you almost nothing about an operation. The spread tells you where the process is loose, who is carrying it, and which cases fall through. Report the distribution and name what sits outside it.
- 04
Systems carry the load
If performance drops when one person is on leave, you do not have a process. You have a dependency. Design so the competent thing happens by default.
- 05
One name per thing
Every queue, standard, number and process has a single owner accountable for it. Groups do not own things. Neither do committees.
- 06
Agents are staff
An AI agent gets a scope, a standard, an evaluation and a retirement date, the same as anyone else. A vendor's resolution rate is not a performance review.
- 07
Intent before execution
State what a piece of work is for before you state how it runs. Most operational disputes are disagreements about purpose being argued as disagreements about process.
- 08
Outlive the reset
Every reorganization, tool migration and leadership change erases whatever was only in somebody's head. Write the operation down in a form that survives the people who built it.
- 09
Always on
Operational truth does not arrive on Monday morning in a deck. The read runs continuously, and the operation should be able to tell you what changed while you were looking elsewhere.
- 10
Decide and keep moving
A decision made this week and corrected next week costs less than a decision debated for a month. Pick the option you can reverse, and write down what would change your mind.
Practices
Principles are what you believe. Practices are what you do on a Tuesday. Adopt them in roughly this order, because each one assumes the one above it.
- 01
Write the competency map before anything else
Not roles, not job titles. The list of things the operation has to be able to do. Everything else attaches to this list, which is why it comes first.
- 02
Write the standard for the person doing the work
Service principles belong in plain language somebody can apply mid-conversation, not in policy prose written for an audit.
- 03
Turn each principle into something you can fail
A principle nobody can violate is decoration. Give each one a named failure condition and a ceiling on the score when it appears.
- 04
Give every queue a stated intent and a named owner
One sentence on what the queue is for. One name against it. Revisit both when the volume or the shape of the work changes.
- 05
Sample continuously
Monthly review of a handful of conversations measures the reviewer. Read enough, often enough, that the number moves before the customer tells you.
- 06
Hold humans and agents to one standard
Run the same evaluation across both. Where an agent fails a criterion a person passes, you have found a scope problem rather than a technology problem.
- 07
Keep the knowledge base small enough to be correct
Every article you cannot maintain is a future wrong answer delivered confidently. Retire aggressively. Two articles that disagree cost more than one that is missing.
- 08
Route by capability
Send work to whoever can actually do it, which requires the competency map to be current. Routing on availability alone spreads failure evenly across the team.
- 09
Measure the outcome, not the vendor's number
Deflection, containment and resolution as reported by a platform are definitions chosen by the seller. Define resolution yourself, in terms of what the customer ended up with, and measure against that.
- 10
Plan capacity from demand
Start from the shape of the work: arrival pattern, handling time by type, the share an agent can hold at standard. Headcount is the output of that model, not the input to it.
- 11
Write the hiring rubric before you meet anyone
Score against the competency map. A rubric written after the interviews is a record of who you liked.
- 12
Put the variance in the report
Lead with the spread and name what sits outside it. A report containing only averages gives nobody anything to act on.
- 13
Keep a decision log
Record what you chose, what you rejected, and what would change the decision. This is the cheapest defense available against the reset.
- 14
Close one loop per cycle, in the open
Pick the finding you will act on, ship the change, publish what moved. An improvement nobody hears about does not change behavior.
What it covers
Agent operations breaks into seven functions. Underneath all of them sits the competency framework, which is the prerequisite rather than a function in its own right. If the map does not exist, nothing above it can be read.
- PerformThe standard, the evaluation, and the daily read of how the work is going.
- BuildProcess design, routing, agent scope, and a record of the decisions behind them.
- EnableOnboarding, training, knowledge. Everything that makes capability transferable.
- MeasureDefinitions, instrumentation, and reporting that names variance instead of hiding it.
- ImproveFindings turned into changes, with owners and dates attached.
- KnowThe intelligence layer. What the work is telling you about the rest of the business.
- GrowCapacity, org design, hiring rubrics, and career paths built from competency.
The four sections that follow are the ladder. Each one describes a state most operations pass through, what is actually wrong at that state, the work that moves you off it, and the test that tells you it worked. Operations are rarely at one level across all seven functions. Start with the lowest rung you are standing on, because the rungs above it will not hold.
Running on people
What this looks like
Work gets done because particular people do it. Your best agent handles the hard cases, and everyone knows who that is. The queue clears because someone stayed late. There is no written definition of good work, so quality is defined after the fact, by which conversations escalated. Documentation exists in fragments: a Slack thread, a spreadsheet one person maintains, a macro library nobody has audited in a year. When leadership asks why a number moved, you go and ask someone.
If you have an AI agent at this stage, it is running close to a vendor default and reporting a resolution rate you did not define. Nobody has read its conversations in bulk.
What is actually wrong
The problem is not effort. Reactive operations are usually full of people working hard and caring a great deal. The problem is that none of that effort accumulates. Every good decision lives inside the person who made it, and when they leave, take a holiday or move teams, the operation loses the capability without anything appearing on a report.
This is also why reactive operations feel unpredictable to everyone outside them. Leadership asks for a forecast and gets a guess, because there is nothing underneath the guess. That erodes the operation's standing long before it erodes its performance, and standing is what you need to fund the fix.
What to do
-
Write the competency map
Not roles. The things the operation has to be able to do, at the level of "identify a carrier exception and set the customer's expectation correctly." Somewhere between ten and thirty entries is normal. Most teams find two or three competencies that nobody actually owns.
-
Write the standard in plain language
A short set of service principles, written for the person handling the conversation. Around eight is workable. Each principle has to be something a reasonable person could apply in the moment, and something they could be shown to have violated.
-
Name owners
Every queue, every principle, every recurring number. One name each. Expect this step to surface the real problem, which is usually that several important things have no owner while one person has eleven.
Do not start measuring yet. Instrumentation on top of an undefined standard produces numbers that get argued about and then ignored, and it burns the credibility you will need later.
Someone who joined last week can be told what good work looks like without asking a senior agent, and you can point at any queue and name its owner without checking.
Getting it written down
What this looks like
The standard exists. There is a document, probably a quality form somebody built in a spreadsheet, and a sample of ten conversations reviewed each month by a person who is also doing three other jobs. The knowledge base has grown faster than anyone can maintain. The AI agent is live and has a dashboard, and the dashboard says something you cannot quite reconcile with what customers are telling you.
Most teams sit here for years. It is a comfortable rung, because from the outside it looks finished. There are documents. There are meetings. There is a monthly report with a chart in it.
What is actually wrong
Written and read are different things. The standard exists but almost nothing is read against it, so it drifts out of contact with the work. Within six months the document describes an operation that no longer exists, everyone knows it, and nobody says so out loud.
The second failure is the join. Quality scores live in one place, satisfaction in another, handling data in a third, knowledge in a fourth. Nothing connects them, so when a measure moves you can describe the movement but not attribute it. You end up explaining outcomes with narrative, and narrative is what gets you overruled in a leadership meeting by someone with a simpler story.
What to do
-
Convert the standard into criteria that can fail
Take each service principle and write the specific behaviors that violate it. Give each one a ceiling, so that when it appears the conversation cannot score above a certain level regardless of tone or outcome. Hallucination. Asking for information the customer already gave. Handing off without saying so. Breaking a stated policy. This is what turns a document into an instrument.
-
Pick the join key and use it everywhere
Competency is the one that holds, because it survives tool migrations and reorganizations and because it sits causally upstream of the outcomes. Tag evaluations, knowledge articles, routing rules and hiring criteria with the same identifier.
-
Move sampling from monthly to continuous
Coverage matters more than depth at this level. Reading a hundred conversations shallowly against explicit criteria beats reading ten deeply against a feeling. Run it across human and agent work on the same criteria from the start, so you never have to reconcile two systems later.
-
Retire knowledge
Delete or merge every article you would not stand behind if an agent quoted it verbatim to a customer. This is unglamorous, it is nobody's favorite quarter, and it is the highest-return work available at this level.
A measure drops and you can say which competency is failing and show the conversations. Not "satisfaction is down four points" but "refund eligibility explanations are failing, here are sixty examples, and the article behind them changed on the ninth."
Reading against the standard
What this looks like
Standards exist and are read against continuously. Ownership is clear. Evaluation runs across people and agents on the same criteria. Reporting shows distribution and names outliers. You can attribute movement to a competency and produce the evidence in an afternoon.
What you have now is evidence, and more of it than the operation can act on. The characteristic failure at this level is a well-instrumented team that reports beautifully and changes slowly. Findings pile up. The improvement backlog grows. Everyone can see the problem and nobody owns the fix.
What is actually wrong
Reading is not acting. A measurement system with no decision rule attached produces a monthly ritual of noticing. It also produces a particular kind of fatigue, because people are being shown their failures on a cadence without being given the means to correct them.
The second problem is scope. Once agents are evaluated properly you will find that a meaningful share of their failures are not model failures. They are scope failures, where the agent was pointed at work that was never defined well enough for anyone to do correctly. Teams at this level usually respond by tuning the agent, which is an expensive way to fix a definition problem.
What to do
-
Attach a decision rule to every measure that matters
When this drops below this level, this named person opens this kind of work by this date. Write it in advance, while nobody is defensive, and treat a triggered rule as information rather than a verdict on a person.
-
Run the improvement backlog like a product backlog
Findings become items. Items get owners and dates. Close one per cycle and publish what moved, internally at minimum. A backlog with no closure rate is a list of grievances.
-
Treat agent scope as a design decision you revisit
Pull work out of an agent's scope when the evaluation says the standard cannot be met there. Put it back when the definition or the knowledge behind it has changed. Record both moves and the reasoning, because the next person to ask will be a procurement conversation.
-
Connect hiring and capacity to the same map
Rubrics score the competencies you already evaluate against. Capacity models start from demand shape and the share of each competency an agent can hold. At this level these stop being separate exercises run by separate people.
A drop in a measure produces a named action with an owner and a date, without anyone convening a meeting to decide whether it should.
Improvement without instruction
What this looks like
The operation improves without being told to. Findings route to owners on their own. The standard itself is under revision based on evidence rather than opinion, and there is a record of why it changed. Capacity, hiring, routing and evaluation all read from the same competency map, so a change in one shows up in the others without anybody rekeying it.
The work here is different in kind. You are no longer building the operation. You are maintaining the conditions under which it keeps building itself, which mostly means protecting the map, the standard and the record from decay.
What is actually wrong
Two things go wrong at the top of the ladder and both are quiet.
The first is capture of the standard. Once evaluation is automated, the standard starts optimizing for what is easy to detect rather than what matters. Criteria that are cheap to score crowd out the ones that are not, and the operation slowly gets very good at the measurable half of its job. Review the standard against actual customer outcomes on a fixed cadence, and be willing to keep a criterion that is expensive to evaluate.
The second is the reset. Optimized operations have the most to lose, because they carry years of accumulated structure and none of it appears on an org chart. A new leader, a platform migration or a reorganization can erase most of it in a quarter, and the erasure stays invisible until performance moves six months later, by which point the cause is unrecoverable. Assume it will happen.
What to do
-
Put the standard on a revision cycle
Named evidence requirements, a fixed cadence, and a written rationale for every change. Changes to the standard deserve the same rigor as changes to the product.
-
Extend the map forward
Workforce planning, org design and career paths built from the same competencies, so progression means demonstrable capability rather than tenure. This is also what makes the map worth maintaining to people outside your function.
-
Keep the record portable
Map, standard, criteria, decision log, evidence. If any of it exists only inside one platform, it is one procurement decision away from gone. Make sure more than one person can explain why each piece exists.
The operation survives the people who built it. Someone can leave and the capability stays.
A working document
This is a description of what has worked, not a specification. Operations differ, and the ladder is a way of sequencing work rather than a grading system. Being reactive in one function and defined in another is the normal case, not a sign of failure.
The parts most likely to change are the practices, because they depend on what tooling makes cheap. The principles have held longer, and the first one has held longest. Everything an operation can do resolves to a competency somebody holds. Build from there.
Maintained by Haven. Version 1.0. If you want to know which rung you are standing on, start the read.