OVERVIEW / THE BUILDING APPROACH

Start with the workload.
Not the framework.

I distilled this theoretical framework and methodology from my own engineering experience.

THE GOAL

Put uncertain model capabilities inside a controllable, reliable, scalable and evaluable production system.

From business logic to validation and review.

CLASSIFY → ARCHITECT → EXTEND → VALIDATE

Four engineering decisions

Choose the execution model. Then place the implementation.

01EXECUTION MODEL

WorkflowPredefined execution path
AgentRuntime reasoning & decisions

02IMPLEMENTATION

Service

Both Workflow and Agent can call services. Encapsulate stable, deterministic steps here.

Choose the form. Select services. Fill the gaps. Check and review.

My engineering methodology · Design examples, not claims that every component is implemented.

01 / CLASSIFY THE WORKLOAD

Not every task
needs an Agent.

Use a Workflow for a predefined path. Use an Agent when the task needs runtime reasoning and dynamic decisions.

Workload determines harness requirements.

FIRST: EXECUTION PATH × DURATION

Three business examples. One map.

Four workload quadrants by execution path and duration. Long Agent describes the category; the other quadrants use the source document’s business examples.
Time
Workflow

Predefined execution path

Agent

Runtime-determined path

LongLONG
Long WorkflowExample

Process & ingest data

A fixed pipeline that runs for a long time. Consider queues, workers and checkpoints.

Long AgentCategory

Dynamic path, long execution

Decide the next step at runtime, with persistence, checkpoints and recovery.

ShortSHORT
Short WorkflowExample

Query & visualize routes

Query → deduplicate / downsample → display. A regular API is enough; no Agent needed.

Short AgentExample

Analyze data & plan collection

The goal is known. Decide whether to query historical routes or statistics as analysis unfolds.

THEN: IS THE USER WAITING?

Interaction mode is a separate dimension.

Foreground

User waits or stays in the loop.
Prioritize first response and streaming.

Background

User need not keep waiting.
Return a task ID; notify on completion.

Classify → derive harness requirements → choose the architecture.

Source: Harness Engineering Methodology · §1.2–1.6

02 / ARCHITECT THE HARNESS

The model reasons.
The harness makes it reliable.

First, the architecture. Then, how each layer is instrumented and governed.

Four layers. Two cross-cutting concerns.

Use this as a reference architecture; adopt only the layers the workload requires.

Four layers. Two cross-cutting concerns.

A / RESPONSIBILITY VIEW
HARNESS · THE PRODUCTION EXECUTION SYSTEMFour layers + two cross-cutting concerns
01

Agent Runtime

EXECUTION

How does the selected Agent execute the task?

Planning & reasoning · tool orchestration · conversation & context
State persistence & resume · runtime limits

02

MCP Gateway

CAPABILITY ACCESS

What can this identity discover and execute?

Discovery & routing · capability exposure · tool authorization
Rate limiting · retry · circuit breaking & failure isolation

03

Custom Services

BUSINESS & EXTENSIONS

Where do business capabilities and behavior extensions live?

MCP Servers Business logic, validation, resource authorization
Middleware Bridge, credential broker, context manager
04

Infrastructure

FOUNDATION

What can run, persist, access and recover?

Compute · network · databases & storage
Quotas · autoscaling · backup · IAM boundaries

A responsibility view, not a detailed request-flow diagram.

How does each layer implement observability and authorization?

B / IMPLEMENTATION DETAIL
ENGINEERING IMPLEMENTATION MATRIXIMPLEMENTATION MATRIX · 4 × 5
Metrics, tracing, logs, audit and authorization implementations across four architectural layers.
LayerOBSERVABILITYSECURITY BOUNDARY
MetricsCollect measurementsTraceConnect the executionLogsRecord and collectAuditRecord sensitive actionsAuthorizationEnforce permissions
01

Agent Runtime

EXECUTION
Metrics

Event / callback → metrics adapter

/metrics → Prometheus scrape

Trace

Create or inherit a trace

Run → LLM → tool → downstream

Logs

stdout / stderr → Fluent Bit → Loki

Correlate trace_id and conversation_id

Audit

Security hooks → audit events

Approvals · denials · tool dispatch

Authorization

Task constraints + scoped capabilities

Constrain actions; do not replace resource authorization

02

MCP Gateway

CAPABILITY ACCESSContextForge · built-in
Metrics

Enable built-in metrics

Requests · latency · errors · rate limits / breakers

Trace

Use native tracing

Accept traceparent → gateway / tool span → downstream

Logs

Enable built-in logs

Access · routing · upstream failures · retries

Audit

Enable built-in audit

Authorization decisions · denials · admin actions

Authorization

RBAC + token scopes

Enforce tool access at tools/call

03

Custom Services

BUSINESS & EXTENSIONS
Metrics

Shared observability SDK

Business counts · dependency latency · queue depth

Trace

OTel auto-instrumentation

Add manual business spans

Logs

Structured JSON logs

Correlate trace_id and request_id

Audit

Sensitive-operation hooks → audit

Mutations · resource access · credential use

Authorization

MCP server: operation + resource checks

Validate business rules; distrust model-supplied identity

04

Infrastructure

FOUNDATION
Metrics

node_exporter / cAdvisor

DB exporters

Trace

Application-side traces

Expose infrastructure dependency latency

Logs

systemd / journal / container logs

DB / proxy logs

Audit

auditd / DB audit / IAM logs

Privilege changes · infrastructure admin actions

Authorization

DB roles / IAM / K8s RBAC

UID/GID · file permissions · network policy

A design view, not a fixed call sequence or a claim that every item is implemented.Gateway-native observability follows the author's current design.
Continue: Custom services →

Source: the author’s Harness Engineering Methodology · §§2–4, §7.2

03 / EXTEND WITH CUSTOM SERVICES

Extend capabilities.
Shape execution.

Extend only where the standard harness falls short. Separate missing capabilities from missing execution controls.

Capability extension ≠ behavior extension.

Design and implement extensions on a case-by-case basis, driven by the workload and its specific requirements.

01

CAPABILITY EXTENSION

MCP Server

Add what the Agent can actually do.

ROLE

Tool implementation · business logic · backend integration
Stable, deterministic processing

BOUNDARY

Argument validation · resource authorization

EXAMPLES

DB / Git / Docs / Kubernetes MCP

02

HARNESS BEHAVIOR EXTENSION

Middleware

Change how the harness controls, governs and optimizes execution.

ROLE

Routing · credential brokerage · context management

METHODS

Policy adapters · observability adapters · caching

EXAMPLES

Bridge / Tool Router / Context Manager

Move stable steps into the service.

AGENT ↔ SERVICE BOUNDARY

Do not make the Agent orchestrate every low-level step

git_checkout→build_image→push_image→update_deployment→wait_rollout→health_check

When this deployment path is stable and deterministic, encapsulate it in a service.

Expose one business tool; let the service own the steps

AgentReason & decide
deploy_service()MCP Server

checkout · build · push · deploy · rollout · verify · rollback

INTENDED BENEFITSFewer tool callsSmaller contextEasier unit / integration testing
Next: validate reliability, scalability and task outcomes.

Source: the author’s Harness Engineering Methodology · §§5–6, §20.2
Deployment is a design example, not a delivery claim.

04 / VALIDATE PRODUCTION QUALITY

Running is not enough.
Validate with evidence.

Reliability, scalability and evaluation: handle failures, accommodate growth and verify outcomes.

A design review framework, not an automated scan.

The Checker is the review checklist used in the Validate step.

Three dimensions. Six checks.

IMPLEMENTATION → EVIDENCE

Swipe horizontally to compare implementation and evidence →

Reliability, scalability and evaluation: six checks, implementation approaches and evidence to review.
Dimension Check Engineering approach Evidence to review
01

Reliability

Recover after failure. Stay usable during it.

A backup is not recovery until restore is tested.

Recoverability

Detect → probe again → restart / replace → restore state → resume / escalate.

Restore control state and workspace / artifacts. Reacquire credentials on resume.

Restore tests
RPO / RTO

Availability

Retry transient failures; make side effects idempotent. Combine degradation, circuit breakers and redundancy.

Make degraded behavior explicit; do not return partial results as normal.

Retry conditions
Idempotency key
Fallback / failover policy

02

Scalability

More tasks. More complex tasks.

Worker pools scale volume; decomposition scales task complexity.

Concurrency

Admission → queue → scheduler → worker pool. Externalize state and bound concurrency.

Scale against queued work and available slots, not CPU alone.

Queue depth / wait
Available worker slots
Active conversations

Complexity

Decompose tasks or use supervisor–executor. Bound parallelism, depth, steps and tokens.

Gateways filter capabilities; services encapsulate stable steps.

Subtask inputs / outputs
Checkpoint / status
Fan-out & runtime limits

03

Evaluation

Did the task succeed? Did the system perform?

Observability provides evidence. Evaluation interprets it.

Task-level

Define success by workload. Prefer tests, schemas and SQL / API checks; use a rubric-based LLM judge for subjective output.

Success criteria
Versioned evaluation cases
Regression results

System-level

Measure responsiveness, end-to-end performance and errors. Use traces to locate bottlenecks.

TTFT · E2E · P50 / P95 / P99
Throughput · error rate
Traces

CLOSE THE LOOP

Every production failure becomes a future regression test.

Real tasks / failuresVersioned eval setCustom evaluatorCI/CD regression gate

Version Agent · Model · Prompt · Dataset

Start with the workload. Close with evidence.

Source: the author’s Harness Engineering Methodology · §§8–27. Review criteria, not claims of implementation or validation.