Agent Runtime
EXECUTIONHow does the selected Agent execute the task?
Planning & reasoning · tool orchestration · conversation & context
State persistence & resume · runtime limits
OVERVIEW / THE BUILDING APPROACH
THE GOAL
Put uncertain model capabilities inside a controllable, reliable, scalable and evaluable production system.
CLASSIFY → ARCHITECT → EXTEND → VALIDATE
Choose the execution form.
Analyze the business logic: Workflow or Agent, short or long-running, foreground or background.
Select the right services.
Use the architecture to select the runtime, gateway, services and infrastructure the workload needs.
Fill the capability gaps.
Add MCP servers or middleware only where the selected services do not meet the requirements.
Check, review and improve.
Use the Checker to assess reliability, scalability and outcomes. Review the findings and refine the design.
01EXECUTION MODEL
02IMPLEMENTATION
ServiceBoth Workflow and Agent can call services. Encapsulate stable, deterministic steps here.
Choose the form. Select services. Fill the gaps. Check and review.
My engineering methodology · Design examples, not claims that every component is implemented.
01 / CLASSIFY THE WORKLOAD
Use a Workflow for a predefined path. Use an Agent when the task needs runtime reasoning and dynamic decisions.
Workload determines harness requirements.Three business examples. One map.
| Time | Workflow Predefined execution path | Agent Runtime-determined path |
|---|---|---|
| LongLONG | Long WorkflowExample Process & ingest dataA fixed pipeline that runs for a long time. Consider queues, workers and checkpoints. |
Long AgentCategory Dynamic path, long executionDecide the next step at runtime, with persistence, checkpoints and recovery. |
| ShortSHORT | Short WorkflowExample Query & visualize routesQuery → deduplicate / downsample → display. A regular API is enough; no Agent needed. |
Short AgentExample Analyze data & plan collectionThe goal is known. Decide whether to query historical routes or statistics as analysis unfolds. |
Interaction mode is a separate dimension.
User waits or stays in the loop.
Prioritize first response and streaming.
User need not keep waiting.
Return a task ID; notify on completion.
Classify → derive harness requirements → choose the architecture.
Source: Harness Engineering Methodology · §1.2–1.6
02 / ARCHITECT THE HARNESS
First, the architecture. Then, how each layer is instrumented and governed.
Four layers. Two cross-cutting concerns.Use this as a reference architecture; adopt only the layers the workload requires.
Planning & reasoning · tool orchestration · conversation & context
State persistence & resume · runtime limits
Discovery & routing · capability exposure · tool authorization
Rate limiting · retry · circuit breaking & failure isolation
Compute · network · databases & storage
Quotas · autoscaling · backup · IAM boundaries
| Layer | OBSERVABILITY | SECURITY BOUNDARY | |||
|---|---|---|---|---|---|
| MetricsCollect measurements | TraceConnect the execution | LogsRecord and collect | AuditRecord sensitive actions | AuthorizationEnforce permissions | |
01Agent RuntimeEXECUTION | Metrics Event / callback → metrics adapter /metrics → Prometheus scrape | Trace Create or inherit a trace Run → LLM → tool → downstream | Logs stdout / stderr → Fluent Bit → Loki Correlate trace_id and conversation_id | Audit Security hooks → audit events Approvals · denials · tool dispatch | Authorization Task constraints + scoped capabilities Constrain actions; do not replace resource authorization |
02MCP GatewayCAPABILITY ACCESSContextForge · built-in | Metrics Enable built-in metrics Requests · latency · errors · rate limits / breakers | Trace Use native tracing Accept traceparent → gateway / tool span → downstream | Logs Enable built-in logs Access · routing · upstream failures · retries | Audit Enable built-in audit Authorization decisions · denials · admin actions | Authorization RBAC + token scopes Enforce tool access at tools/call |
03Custom ServicesBUSINESS & EXTENSIONS | Metrics Shared observability SDK Business counts · dependency latency · queue depth | Trace OTel auto-instrumentation Add manual business spans | Logs Structured JSON logs Correlate trace_id and request_id | Audit Sensitive-operation hooks → audit Mutations · resource access · credential use | Authorization MCP server: operation + resource checks Validate business rules; distrust model-supplied identity |
04InfrastructureFOUNDATION | Metrics node_exporter / cAdvisor DB exporters | Trace Application-side traces Expose infrastructure dependency latency | Logs systemd / journal / container logs DB / proxy logs | Audit auditd / DB audit / IAM logs Privilege changes · infrastructure admin actions | Authorization DB roles / IAM / K8s RBAC UID/GID · file permissions · network policy |
Source: the author’s Harness Engineering Methodology · §§2–4, §7.2
03 / EXTEND WITH CUSTOM SERVICES
Extend only where the standard harness falls short. Separate missing capabilities from missing execution controls.
Capability extension ≠ behavior extension.Design and implement extensions on a case-by-case basis, driven by the workload and its specific requirements.
CAPABILITY EXTENSION
Add what the Agent can actually do.
Tool implementation · business logic · backend integration
Stable, deterministic processing
Argument validation · resource authorization
DB / Git / Docs / Kubernetes MCP
HARNESS BEHAVIOR EXTENSION
Change how the harness controls, governs and optimizes execution.
Routing · credential brokerage · context management
Policy adapters · observability adapters · caching
Bridge / Tool Router / Context Manager
Do not make the Agent orchestrate every low-level step
When this deployment path is stable and deterministic, encapsulate it in a service.
Expose one business tool; let the service own the steps
checkout · build · push · deploy · rollout · verify · rollback
Source: the author’s Harness Engineering Methodology · §§5–6, §20.2
Deployment is a design example, not a delivery claim.
04 / VALIDATE PRODUCTION QUALITY
Reliability, scalability and evaluation: handle failures, accommodate growth and verify outcomes.
A design review framework, not an automated scan.The Checker is the review checklist used in the Validate step.
IMPLEMENTATION → EVIDENCE
Swipe horizontally to compare implementation and evidence →
| Dimension | Check | Engineering approach | Evidence to review |
|---|---|---|---|
01
ReliabilityRecover after failure. Stay usable during it. A backup is not recovery until restore is tested. |
Recoverability | Detect → probe again → restart / replace → restore state → resume / escalate. Restore control state and workspace / artifacts. Reacquire credentials on resume. |
Restore tests |
| Availability | Retry transient failures; make side effects idempotent. Combine degradation, circuit breakers and redundancy. Make degraded behavior explicit; do not return partial results as normal. |
Retry conditions |
|
02
ScalabilityMore tasks. More complex tasks. Worker pools scale volume; decomposition scales task complexity. |
Concurrency | Admission → queue → scheduler → worker pool. Externalize state and bound concurrency. Scale against queued work and available slots, not CPU alone. |
Queue depth / wait |
| Complexity | Decompose tasks or use supervisor–executor. Bound parallelism, depth, steps and tokens. Gateways filter capabilities; services encapsulate stable steps. |
Subtask inputs / outputs |
|
03
EvaluationDid the task succeed? Did the system perform? Observability provides evidence. Evaluation interprets it. |
Task-level | Define success by workload. Prefer tests, schemas and SQL / API checks; use a rubric-based LLM judge for subjective output. |
Success criteria |
| System-level | Measure responsiveness, end-to-end performance and errors. Use traces to locate bottlenecks. |
TTFT · E2E · P50 / P95 / P99 |
CLOSE THE LOOP
Version Agent · Model · Prompt · Dataset
Start with the workload. Close with evidence.
Source: the author’s Harness Engineering Methodology · §§8–27. Review criteria, not claims of implementation or validation.