Vīkṣya › Insights

Picking the First AI-Driven Delivery Pilot — Why the Project You Choose Decides Whether the Model Spreads

October 6, 2026

Picking the First AI-Driven Delivery Pilot — Why the Project You Choose Decides Whether the Model Spreads

Executive Summary. The advice most organizations get on their first AI-driven delivery pilot is advice about AI use cases — which problem is valuable enough, and tractable enough, to build an AI system for. That's a real question. It is not the question an organization faces the moment it decides to run its delivery process itself, not just a product, on an AI-driven model for the first time. That decision is "which existing piece of work do we hand to this new way of working, and what happens to the rest of the transition if that piece goes badly." Few organizations treat this as a decision with its own criteria. Most default to whatever is next on the roadmap, or whatever team volunteers first. This piece lays out six concrete questions to ask instead, a simple way to score the answers, and a worked comparison across three realistic candidates — the method we use when we help a client choose theirs.

The Question That Isn't About the Technology

Ask ten engineering leaders how they picked their organization's first AI-driven delivery pilot, and most answers will be a reason dressed up after the fact: it was next on the roadmap, a team volunteered, or a sponsor wanted a specific project accelerated. None of those answers are dishonest. They're just answers to the wrong question. The question that actually matters is narrower: if the new delivery model works on this project, does the rest of the organization have a reason to believe it will work on theirs too? And if it doesn't work, will the explanation be specific to this project, or will it read as a verdict on the whole approach?

That distinction gets collapsed constantly, and collapsing it is expensive. A pilot that stalls for reasons that have nothing to do with the delivery model — requirements that were never clear, a team already stretched across three other priorities, a dependency on a system two other teams are mid-migration on — still lands, to everyone watching from outside, as "we tried the AI-driven approach and it didn't work." The delivery model never gets a fair trial. It gets blamed for conditions it never controlled, and the organization's appetite for a second attempt drops accordingly.

Two Extremes, One Shared Mistake

There's a natural pull toward one of two choices when a pilot gets picked, and both fail for the same underlying reason: neither actually produces a result anyone should trust.

The first is the safe choice — small scope, no real stakeholders paying attention, low consequence if the project slips. The instinct behind it is reasonable: nobody wants the public debut of a new way of working to be the one that breaks something important. But a project with no real constraints doesn't exercise the parts of a delivery model that matter under pressure: how it handles requirements that arrive half-formed, how review scales when the output actually matters to someone, how a team recovers mid-sprint when something goes sideways. A pilot that never meets friction never demonstrates it can survive friction. It proves the model works when nothing was ever at risk, which nobody was doubting.

The second is the ambitious choice — the highest-visibility project, picked specifically to make the loudest possible statement if it succeeds. This is where most stalled transitions actually start. A high-stakes project arrives with its own baggage: an external deadline nobody can move, a customer-facing surface where a visible stumble has a real cost, dependencies outside the team's control. When that kind of pilot runs into trouble, nobody can separate the delivery model from the project's pre-existing difficulty. The organization doesn't learn whether AI-driven delivery works. It learns that a hard project was hard, which it already knew.

A Practical Way to Score the Choice

Six questions, asked of every candidate project, do most of the real work. None of them concern the AI system or product the pilot team happens to be building — they concern whether the project itself, as a test instrument, is actually built to produce a trustworthy answer.

  1. Is the scope real but bounded? Enough genuine complexity that the team has to exercise judgment on real decisions, not so much that a failure has no clear, attributable cause.
  2. Does the team actually want to run it this way? A pilot handed to a reluctant team tests morale, not the delivery model. Willingness isn't a courtesy to extend — it's a variable you're choosing to introduce or avoid.
  3. Is there a pre-agreed, specific way to know it worked? Not "did it feel smoother," but a stated measure, written down before the pilot starts, not assembled afterward to match whatever happened.
  4. Can the project absorb a stumble without it becoming a crisis? The gap between "this slip cost a week" and "this slip is now an executive escalation" decides whether a rough patch becomes a lesson or gets hidden.
  5. Would a skeptical peer team recognize themselves in it? A pilot that looks nothing like the work most other teams actually do won't travel if it succeeds — "it worked there" needs to plausibly mean "it could work here."
  6. Does the timeline have genuine slack? Not an open-ended schedule — a pilot still needs a deadline to be a real test — but enough margin that one bad week doesn't cascade into a missed external commitment.

A simple way to use these six questions without a scoring tool: rate each candidate project 1 (fails the question), 2 (partially meets it), or 3 (clearly meets it), for a possible total of 6 to 18. In our own work, a project scoring 14 or above is a strong pilot candidate; one scoring below 10 usually fails on either question 1 (too easy, nothing at stake) or question 4 (too fragile, one bad week becomes a crisis). The number matters less than the exercise of writing down a reason for every score — a 2 with no justification is a guess wearing a number.

A Worked Comparison

Picture three candidate projects on the same roadmap, each a realistic stand-in for what most engineering organizations actually have waiting.

Project A: an internal reporting dashboard. Real, useful work, but low visibility and low consequence — a one-week slip inconveniences a handful of analysts. Scope: real but narrow (score 2). Team willingness: high, nobody minds (3). Success measure: vague — "looks right to the finance team" (1). Absorbs a stumble: trivially (3). Peer recognition: low — most teams don't build internal dashboards (1). Timeline slack: generous, no external pressure (3). Total: 13. Close to a strong candidate on paper, but the missing success measure and weak peer recognition are the real problem — a result here will be hard to defend as evidence of anything.

Project B: a customer-facing billing migration. High visibility, real consequence, and a hard deadline tied to a vendor contract. Scope: real and substantial (3). Team willingness: mixed — some engineers are uneasy about a new delivery model on anything customer-facing (2). Success measure: clear and already tracked (billing accuracy, migration completeness) (3). Absorbs a stumble: poorly — the vendor deadline turns any slip into an escalation (1). Peer recognition: high, most teams have run something like it (3). Timeline slack: almost none (1). Total: 13. The same score as Project A, for the opposite reason: this project would produce a genuinely informative result if it succeeded, but a failure of any kind — delivery-model-caused or not — becomes a vendor-facing incident rather than a contained lesson.

Project C: an internal claims-triage workflow rebuild. Moderate visibility inside operations, moderate consequence, no external contract. Scope: real and bounded — a defined set of triage rules, a known data set (3). Team willingness: volunteered, genuinely curious about the new model (3). Success measure: agreed in advance — triage accuracy and processing time against the current baseline (3). Absorbs a stumble: a slip delays an internal rollout by a sprint, no outside party notices (2). Peer recognition: high — most operations-adjacent teams run comparable workflow projects (3). Timeline slack: two sprints of margin built in from the start (2). Total: 16.

Project C is the pilot worth running, and the scoring makes the reason visible rather than a matter of taste: it's the only candidate that clears every question without a glaring weak point, not the only candidate with a high total. A's total looks similar but hides a missing success measure. B's total hides a deadline that converts any stumble into a crisis. The total is a starting conversation, not a verdict — the column that's weak is usually more informative than the sum.

The Second Question Nobody Asks: What Happens to the Rest of the Backlog

Choosing the first pilot well solves half the problem. The other half is what it implies for everything waiting behind it. An organization that picks one good pilot and has no answer for "which project is next, and why" has made one good decision and left the following fifty to instinct again. The same six questions that chose the first pilot start a segmentation exercise for the rest of the portfolio: which projects are strong second and third candidates, which should wait until the model has matured past its first proof point, and which genuinely don't fit this delivery model at all and shouldn't be forced into it. Skip that step and a successful pilot quietly becomes a one-off — a good story about one team, with no repeatable logic for choosing the next one.

Where This Sits Next to How You Already Prioritize AI Use Cases

This isn't a competing framework to how an organization decides which AI system to build. Business-value-versus-complexity scoring and start-boring guidance both still matter, and neither should be stretched to answer this question too — they weren't built to. Those frameworks ask which problem is worth solving with AI. This one asks which existing piece of delivery work is the right proving ground for a new way of doing the work, independent of what gets built on it. An organization can answer the first question well and still get this one wrong — the AI use case ships fine, while the delivery model it shipped under never gets a fair test, and nobody notices the second failure because the first success is distracting everyone from it.

That's the gap our own AI-DLC Pilot Selection & Portfolio Segmentation Tool was built to close: scoring candidate delivery projects against criteria like the six above to produce a ranked shortlist, and segmenting the rest of the portfolio so the choice after the first pilot isn't back to instinct either.

What to Do Before Your Next Planning Meeting

If a first AI-driven delivery pilot is being chosen in the next few weeks, four things are worth doing before the decision gets made in a room:

  • Write the six questions down and score every real candidate, not just the one or two already being discussed informally. A project that nobody has mentioned often scores better than the obvious pick.
  • Ask every candidate's team directly whether they want it, not whether they'd accept it if assigned. The difference shows up in the pilot's first bad week.
  • Agree the success measure in writing before the pilot starts, and store it somewhere the whole team can see it — not retrofit it from whatever the pilot happens to produce.
  • Name, out loud, what happens if it stumbles. If the honest answer involves a vendor contract, a board deadline, or a customer-facing outage, that project is not this pilot, whatever else recommends it.

The Question Worth Taking Into the Room

The project handed to a first AI-driven delivery pilot is not a scheduling decision. It's the test instrument for the entire transition, and a badly chosen instrument produces a result nobody should trust, whichever way it comes out. Before the next candidate gets proposed, ask whether it would actually tell the organization something — or whether it's already been picked for reasons that have nothing to do with what anyone is trying to learn.