Context Engineering Gave Your Agents a Service Catalog. It Didn't Give Them a Definition of Done.

Executive Summary. The current wave of context engineering work — service catalogs, ownership records, runbooks, incident feeds — solves a real problem: an agent that doesn't know which service it's touching, who owns it, or whether it's healthy will make bad calls. But we keep seeing teams finish that work and still ship bad agent output, because a service catalog answers "what does the agent know about the system" and never touches a separate question: "what does the agent know about the task." An agent can have complete visibility into ownership, health, and deployment history for a service, and still get a specific unit of work wrong, because nobody defined what that task was actually supposed to accomplish, what was out of scope, or what "done" looked like before the agent started. That's a different failure mode, and a service catalog doesn't fix it.
Look at how the term is actually being used right now. Across the recent wave of writing on context engineering, the examples cluster around one layer: does the agent know the service's owner, its repository, its on-call status, its incident history, its deployment record. That's infrastructure context — the facts that are true about a system regardless of which task anyone is currently asking an agent to perform on it. It's genuinely useful, and it's also the easier problem to solve, because it only has to be built once. A platform team stands up a service catalog, connects it to the tools that already track ownership and health, and every future task against that service inherits the same context for free.
Task context doesn't work that way. It can't be centralized once and reused, because it's specific to the unit of work someone is asking for right now. "Fix the checkout bug" is not a task an infrastructure layer can describe for you, no matter how complete the service catalog is underneath it. Does "fix" mean patch the immediate symptom, or trace and fix the underlying cause? Is backward compatibility with the existing API contract required, or is this a breaking change the team already agreed to accept? Should the fix ship with new test coverage, or is a hot-fix on a deadline acceptable without it? An agent with perfect knowledge of the checkout service's ownership, health, and deployment history has no way to answer any of those questions, because none of them are facts about the service. They're facts about the task, and somebody has to state them before the agent starts, not discover after the fact that the agent guessed wrong.
This is why the failure mode looks different from an infrastructure gap. When an agent lacks service context, it fails loudly and early — it can't find the repository, can't resolve the owner, stalls out asking for information it has no way to retrieve. When an agent lacks task context, it doesn't stall. It produces something. The code runs, the pull request opens, the ticket moves to "done." The problem only surfaces later, when a reviewer discovers the fix touched three files nobody expected it to touch, or shipped without the test coverage the team assumed was implied, or solved a more general version of the problem than anyone asked for and introduced risk nobody signed off on. That's not a model capability problem. It's the same failure you'd get handing a competent engineer a one-line ticket and no conversation about what "done" means, then being surprised when their idea of done doesn't match yours.
Stay with the checkout-bug example for a moment, because the shape of the failure matters more than the specific case. Say the agent has full infrastructure context: it knows the checkout service's owner, that it's healthy, that its last deployment was clean, that the on-call rotation is staffed. It reads the bug report, traces the fault to a race condition in how two payment providers' callbacks are handled, and produces a fix. The fix works — in the sense that the race condition is gone. But the team that filed the ticket wanted a minimal patch they could ship same-day, because a promotional campaign was launching that evening and they had already frozen everything else about the payment flow. What they got instead was a fix that also refactored the callback-handling module, because nothing told the agent that a same-day, minimal-surface-area patch was the actual requirement rather than a well-engineered fix. The service catalog was completely accurate. It was also silent on the one thing that determined whether the output was usable that day.
That gap gets worse, not better, as more of the work moves to agents, because the missing step was never really about the agent — it was about whoever used to absorb it informally. An engineer picking up that same ticket would likely have pinged the requester, asked what "fix" meant given the campaign timeline, and adjusted scope before writing a line of code. That back-and-forth was never written down anywhere; it happened in a Slack thread or a hallway conversation, and the informal judgment call covered for the missing specification. An agent doesn't have that hallway. It has whatever was in the prompt, plus whatever the infrastructure layer can tell it about the system — and neither one contains a campaign launch date. Scaling agent use without also scaling how task context gets specified doesn't remove that judgment call. It just removes the person who used to make it by default.
The reason task context keeps getting skipped isn't that anyone thinks it doesn't matter. It's that, unlike a service catalog, it can't be delegated to a platform team and built once. Somebody has to assemble it for every task — and unless that assembly is treated as its own deliberate step, with its own checklist, it quietly collapses into whatever fits in a one-line prompt. Four things belong in that checklist, and they're worth naming precisely because "give the agent more context" is too vague to act on: intent — what this specific task has to accomplish, distinct from the general goal it sits under; scope boundary — what the task must not do, stated explicitly, because silence is not a boundary an agent can infer; constraints — the format, technical, or voice rules the output has to satisfy; and definition of done — concrete enough that someone with no memory of the original request could check the output against it and get the same answer you would.
We think of this as diagnosis work, distinct from the execution work of actually doing the task — the same distinction that shows up across how we think about AI-DLC generally. Our free context-pack-builder skill exists for exactly this diagnosis step: assembling that four-part pack — intent, scope, constraints, definition of done — for one discrete unit of work before any drafting or coding starts, rather than leaving it to whoever's prompt happens to get written that day. For a team doing this often enough that it needs to be a repeatable practice rather than a one-off habit, our context-engineering playbook goes further, defining the fuller set of knowledge artifacts a structured delivery workflow depends on and the patterns for extracting them consistently. Neither replaces the infrastructure layer everyone else is currently writing about. They sit one level below it, answering a question a service catalog was never built to answer.
The test is simple enough to run against your own agent workflows this week. Take the last task you handed to an agent that came back wrong in a way that surprised you. Ask whether the agent was missing information about the system — which is a service-catalog gap — or information about what that specific task was supposed to accomplish, what it should have left alone, and what "done" would have looked like — which is a task gap, and no amount of service-catalog completeness would have closed it. Most teams that have already invested in the infrastructure layer will find it's the second one.