Prompts are important, but they are not an operating model. As soon as an AI workflow uses multiple models, tools, validators and human decisions, reliability becomes a software architecture problem. The system needs explicit contracts at every boundary.
The failure mode: invisible policy
A simplistic agent often hides critical choices inside framework defaults: which model receives the request, how many times a failed call is retried, whether a fallback is allowed and what happens after a tool produces an invalid result. That may be acceptable in a prototype. In production, it makes cost, latency and risk unpredictable.
The reliable unit is not the model call. It is the complete, observable execution path.
Make execution policy a first-class contract
In the Work Harness I am building, each retryable step declares its own attempts, time budget, retryable failure categories, fallback routes and side-effect classification. Unsafe side effects cannot be retried. Retried operations must have both per-attempt and total deadlines.
class StepExecutionPolicy:
max_attempts: int = 1
attempt_timeout: float | None
total_deadline: float | None
retryable: tuple[FailureCategory, ...]
side_effect: "none | idempotent | unsafe"
def validate(self):
if self.max_attempts > 1:
require(self.attempt_timeout)
require(self.total_deadline)
reject(self.side_effect == "unsafe")
This turns behavior that is usually implicit into something reviewable and testable. Authentication or configuration errors fail quickly; transient network and rate-limit failures may retry; output repair follows a separate path rather than being confused with provider availability.
Record facts, not private reasoning
Every execution emits versioned events for routing, stages, model attempts, tools, retries, validation, artifacts and human decisions. Events share stable identifiers and ordering. They record observable facts—route, duration, outcome, usage, sanitized error category and artifact references—not private chain-of-thought.
Separate the public capability from internal specialists
A user should choose a capability—such as software architecture or options analysis—not a particular underlying model. Internally, a graph can route between specialists, providers and reasoning levels according to risk and complexity. This keeps the external contract stable while the implementation evolves.
What this makes possible
Once execution is explicit, the platform can answer questions that prompt logs cannot: Which paths succeed without human intervention? Which fallback increases cost without improving quality? Which validator catches regressions? Where can a deterministic component replace a model safely?
That is the transition from experimenting with AI to engineering an AI system: not removing uncertainty, but making it bounded, observable and governable.