Nobody in the room is asking whether to adopt AI. That question got settled somewhere in the last eighteen months and arguing it now is a way to avoid doing anything.

The question people actually ask, usually about forty minutes into a call, is narrower and much harder: I have twelve things that annoy me, which one do I point this at first.

Here is how I would pick.

The heuristic

Three things, multiplied together. Frequency, annoyance, and how little judgment the step needs.

How often does it happen. Not how painful it is when it happens. How often. A thing that happens four hundred times a month is worth attention even if each instance is mild, because the payoff compounds and, more importantly, you get four hundred chances to find out whether it works. A thing that happens twice a quarter gives you two data points a quarter.

How much does it annoy somebody. This is not a soft criterion. Annoyance is a decent proxy for the work being genuinely unrewarding, and unrewarding work is the work that gets skipped when the week gets bad. It also determines whether anyone will actually adopt the fix.

How much judgment does it need. Start low. Not because judgment work is off limits, but because you cannot evaluate what you cannot check. If a person can glance at the output and immediately say right or wrong, you will know inside a fortnight. If checking it takes forty minutes of expertise, you will not check it, and you will end up trusting something you never verified.

Where to point your first buildAgency tasks plotted by how often they happen against how much judgment they need. Posting payments from the carrier report and reconciling signed documents sit in the high frequency, low judgment corner, which is where to start. Coverage advice and claims conversations sit high on judgment and are not first builds.START HEREHow often it happensRarelyConstantlyJudgment it needsA lotNonePost payments from thecarrier reportReconcile signeddocuments to filesFlag what changed on arenewalMonth end, what isstuck in bindingAdvise on a coveragegapHandle a claimconversation
The shaded corner is not the most impressive work. It is the work you can verify, which is what makes it first.

What that looks like when it works

Three of the plotted examples are real, and it is worth saying what happened with each because the pattern is the argument.

An agency in Mississippi pointed it at the non pay and late pay pipeline. High frequency, near zero judgment, and enormously tedious. Every morning it pulls the carrier report, posts cancellations and reinstatements, and enters payment amounts. Nobody touches it. That work is checkable at a glance: either the payment amount matches the report or it does not.

The same agency ran a reconciliation of signed e-signature documents against customer files. First run surfaced eighty eight signed documents that had never made it onto a file. High volume, low judgment, and the output is a list you can spot check in ten minutes.

Now the interesting one, because it breaks the rule. A renewal comparison caught a carrier quietly moving roof settlement from replacement cost to actual cash value, including on newer homes. Twenty minutes produced a complete list of every affected policy, and the carrier did not have that list.

That is higher judgment than the first two, and it produced the biggest single result of the three. So why is it not the recommended first build? Because the agency that ran it had already done the first two. They knew what the output looked like when it was right, they knew where it tended to be wrong, and they had a habit of checking. The renewal comparison was their third build, not their first, and that sequencing is the whole reason it worked.

The order to work through your list inFirst list what actually happens often. Then cross off anything you could not verify at a glance. Then pick the most annoying one that survives, and run it for two weeks before adding anything else.List what actually happens often, by volumeCross off anything you could not check at aglancePick the most annoying one leftRun it two weeks before you build anythingelse
The last step is the one people skip, and skipping it is how you end up with six half trusted builds.

What not to start with

Anything you cannot check. If verifying the output requires the same expertise as doing the work, you have not saved the work. You have moved it and added a step.

The thing that would impress people. There is always a build that would be a great story. It is a bad first build, because a bad first build teaches your team that this stuff does not work, and that lesson takes a year to unteach.

Anything on the far side of the line. It never binds coverage, it never takes a payment, it never closes a cancellation or a non renewal on its own. Those are not first builds and they are not tenth builds either.

One more filter

Before any of it: can you describe the workflow in plain language, end to end, including who does each step and what causes the next one?

If not, that is your first project, and it is not an AI project. It is the mapping exercise, and it is worth doing even if you never automate the thing, because you cannot automate what you cannot explain.