← All insights
AI Strategy

How to Pick the One Task to Automate First

By Dr. Matt Goodwin  ·  August 18, 2026  ·  10 min read

The hardest part of starting with AI is not the technology. It is choosing where to start. Pick wrong and you spend real effort automating something that did not matter, and worse, you teach yourself and your team that automation does not pay. Pick right and the first win funds and motivates everything after it. This piece turns that choice into an instrument: where candidates come from, the three filters in depth, a scoring protocol with anchored scales and a worked example, the four anti-patterns that waste first efforts, and what to do the moment the winner is chosen.

Where the candidates come from

A prioritization method is only as good as the list it prioritizes, and most owners build the list from memory, which means they build it from whatever annoyed them most recently. Build it from observation instead. The one week logging exercise from the communication tax your business pays every day produces exactly the right raw material: a week of real incidents where a person had to move information by hand, sorted into translation, chasing, drops, and rebuilding, each with a minutes count attached. Run the exercise faithfully and you will finish the week with a candidate list grounded in what your business does rather than what it complains about, and the difference between those two lists is usually the difference between a first automation that pays and one that decorates. Owners have a second instrument pointed at the same target: the personal inventory from the owner’s trap, one week of everything you personally touch, split into judgment work and movement work. The movement list and the incident log converge on the same candidates from two directions, and where they agree, you can trust the finding.

The three filters

Every candidate on the list gets run through three filters, and each filter deserves a real explanation, because the reasoning is what lets you apply them to cases the examples do not cover.

Frequency asks how often the task happens, and it matters because automation is a fixed cost amortized across repetitions. A task you do fifty times a week pays back its build fifty times a week. A task you do twice a month, however irritating, pays back twice a month, and at that rate almost no build is worth its cost. Frequency is also where instinct misleads most reliably: rare tasks feel bigger precisely because they are novel each time, while the truly frequent ones have worn a groove so deep you no longer feel them. Trust the log, not the feeling. The arithmetic behind the filter is worth internalizing: a task’s annual repetition count is the multiplier on every minute the automation saves. Eight times a day is roughly nineteen hundred repetitions a year. Twice a month is twenty four. That is an eighty-fold difference in payback before quality, consistency, or after-hours coverage enter the comparison at all.

Rules asks how predictable the task is, and it measures buildability. Work that follows clear, repeatable steps, when this arrives, do that, then send this, automates cleanly and behaves reliably once automated. Work that requires judgment on every instance does not automate cleanly, and attempting it first buys you an unreliable system in your most visible spot. The filter is a spectrum, not a switch. Plenty of tasks are rule-based for ninety percent of instances with a judgment call in the remainder, and those are fine later, built with a human handling the exceptions. The pattern even has a name worth knowing, human in the loop: the system carries every standard instance and routes the genuine judgment calls to a person, a design choice rather than a compromise. It belongs in your second or third build, not your first. For the first build, though, you want the clean end of the spectrum, because the first build is also the demonstration that convinces everyone, including you, that this works.

Pain asks what it costs when the task goes wrong or gets skipped, and it is the filter that separates expensive problems from loud ones. A dropped follow-up that loses a customer is priced in revenue. A tedious report that bores you is priced in sighs. The distinction maps directly onto the tax categories: drops carry revenue prices, which is why a failure-prone handoff outranks a merely boring chore every time, even when the chore is more annoying to live with.

The scoring protocol

Score every candidate from one to five on each filter, using anchored scales rather than gut feel, because anchors are what make two people’s scores comparable. For frequency: a five is many times daily, a three is a few times weekly, a one is monthly or rarer. For rules: a five is identical steps every single time, a three is standard steps with occasional exceptions, a one is fresh judgment every instance. For pain: a five is failure that costs customers or revenue directly, a three is failure that costs real rework hours, a one is failure that costs nothing but mild embarrassment. Add the three columns. Highest total wins.

A worked example makes the mechanics honest. Take three candidates from a typical service business log. Candidate one, booking confirmations and reminders: frequency five, happens with every job; rules five, identical steps each time; pain four, a missed confirmation produces no-shows, which are priced in revenue. Total fourteen. Candidate two, writing custom proposals: frequency three, weekly; rules two, every proposal is a judgment product; pain three, a bad one costs a deal but a slow one mostly costs hours. Total eight. Candidate three, the monthly bookkeeping tidy-up: frequency one; rules four, mostly mechanical; pain two, failure means an annoying catch-up session. Total seven. Instinct, for the record, usually picks the proposals, because proposals are the task the owner personally dreads. The math picks the confirmations, and the math is right: it is the candidate that runs constantly, automates cleanly, and bleeds revenue when it fails.

The losing candidates are informative too, which is why the scores are worth keeping rather than discarding. The proposals scored low on rules, which does not mean never; it means later, and differently, as assisted drafting with your judgment retained rather than as automation, once the clean wins have built trust in the approach. The bookkeeping scored low on frequency, which suggests the cheapest fix is not a build at all but a calendar: batch it monthly, protect the slot, and spend the automation budget where the meter runs daily.

The four anti-patterns

The filters also name the classic ways first efforts die, and each anti-pattern is a filter ignored.

Automating the annoyance is the frequency failure: picking the rare, hated task because hatred is vivid, then discovering the build cost can never amortize across six occurrences a year. Automating judgment first is the rules failure: pointing the first build at the most human task on the list, then watching it wobble publicly and poison the appetite for everything after. Automating a broken process is the subtler third failure: automation multiplies whatever it touches, including dysfunction, so a confused workflow automated becomes confusion at machine speed. If the log shows a task whose steps nobody can state, the task needs a redesign before it needs a robot. And boiling the ocean is the failure of skipping selection entirely, launching five automations at once, and this one the research prices directly: BCG found that the companies generating real value from AI concentrate on a few high-priority opportunities and pursue roughly half as many initiatives as their struggling peers, going deeper on each instead of spreading thin (BCG, 2024). Choosing one task well is that same discipline at the scale of a single business.

Score the list, keep the list

The scoring produces more than a winner. Sort the whole list by total, descending, and you are holding your build sequence for the next two quarters, already prioritized by leverage instead of by whoever complained loudest in a given week. That standing list is worth maintaining as a living document, because it converts every future what-should-we-automate conversation from a debate into a lookup.

It does have to stay living, though, because the landscape moves when you build. Removing one task changes the scores of others: the chasing that existed to cover a failure-prone handoff disappears when the handoff stops failing, and a task that scored a four on pain can drop to a two once the workflow upstream of it is reliable. So rescore on a cadence, after every completed build or quarterly, whichever comes first. Ten minutes of rescoring is what keeps the second and third builds landing as well as the first one did, instead of aiming at a map of a business that no longer exists.

Why a method beats instinct at all

It is worth being explicit about why the ceremony of scoring earns its ten minutes. My doctoral research found that the leaders who successfully adopted new technology were the ones who started from a concrete, recognized problem, while the ones who stalled had no named problem anchoring the effort (Goodwin, 2014). The scoring protocol is a machine for manufacturing exactly that anchor: it forces the problem to be named, sized, and compared before anything is bought or built. It is the problem-first sequence from AI adoption starts with a problem, not a tool, operationalized into a form that survives a busy week, and it is one of the patterns in the 5 drivers of AI adoption.

The objections

The first objection: what if two candidates tie? Break ties on pain, because failure costs compound while minutes merely accumulate. A task that loses customers when it fails is buying you back revenue and reputation at once, and the second of those keeps paying after the build is forgotten.

The second objection: what if the winner needs integrations we do not have? Then the scoring has done its second job, which is diagnosis. Discovering that your highest-leverage task crosses two systems that do not talk is not a blocker; it is the finding, and the connection is the thing to build. That gap between tools is precisely where the most expensive incidents in the log have been living, which is the argument of the communication tax.

The third objection: should the team weigh in, or is this an owner call? Both, in sequence. The logging week is team-wide, because the incidents live in everyone’s day and the owner sees a fraction of them. The scoring is where the owner’s vantage matters, because pain in particular is priced in revenue and customer consequence, and that is the view from the P&L, not from any single desk. Collect wide, score from the top.

From pick to pilot

The chosen task deserves one more discipline before anything gets built: a boundary. Give it a container with a fixed clock, a defined success measure, and a small blast radius, which is exactly the shape of a well-run first pilot, thirty days, one task, one number that will move if it works. That container is what turns your first automation from an open-ended project into an answered question, and the full construction is in how to run your first AI pilot in 30 days without wasting a quarter.

The move

Run the logging week if you have not, or pull five candidates from memory if you must start today. Score them one to five on frequency, rules, and pain using the anchors above, total the columns, and commit to the winner, not the one you would have picked by feel. It is rarely the same task, and that is the entire point. One permission slip belongs in the method too. Some tasks will score low and stay manual on purpose, because they are your craft, or because you simply like doing them, and that is a legitimate outcome. The instrument prioritizes what to automate. It does not mandate automating everything, and a deliberate keep, chosen with the score in front of you, is a decision rather than a failure. What the method removes is only the accidental keep, the expensive task that stayed manual because nobody ever put a number on it. To pressure-test your pick against where the leverage sits in your operation, the Omnine AI Readiness Assessment takes about three minutes.

References

Goodwin, M. R. (2014). A qualitative descriptive multiple-case study: Fortune 500 leaders’ social business platform adoption (Doctoral dissertation, University of Phoenix). ProQuest Dissertations Publishing (UMI No. 3648813).

Boston Consulting Group. (2024). Where’s the value in AI?

Where do you actually stand with AI?

The Omnine AI Readiness Assessment scores you in about six minutes and shows you exactly where to start.

Take the Assessment →