Retour au blog
The Minimum Viable Agent Platform
AI Agents Publié February 23, 2026 Par Sami Kalliokoski

The Minimum Viable Agent Platform

What You Actually Need Before Calling It Production

There is a lot of noise around AI agents.

Most discussions focus on prompting techniques, tool calling frameworks, or multi agent orchestration patterns. Those are interesting topics. But none of them answer the real production question:

Can this system run safely, predictably, and repeatedly under load?

If not, you do not have an agent platform. You have a demo.

This post describes what I consider the Minimum Viable Agent Platform. Not the most advanced. Not the most academic. The minimum set of architectural capabilities required before an agent system deserves to be called production ready.


First Principle: Agents Are Distributed Systems With a Stochastic Core

An agent system is not just an AI problem.

It is a distributed system that happens to contain a probabilistic decision engine.

That means:

  • It must tolerate partial failures
  • It must enforce boundaries
  • It must expose observability
  • It must control blast radius
  • It must be debuggable

If you ignore these, your system will eventually hurt you. Not because the model is bad, but because the architecture is incomplete.


The 7 Requirements of a Minimum Viable Agent Platform

Without these, you are not production ready.


1. Budget Per Run

Every agent mission must have a predefined budget.

Budget can mean:

  • Token budget
  • Monetary budget
  • Time budget
  • Tool call budget

The important part is not the unit. It is the hard limit.

An agent without a budget is an infinite loop waiting to happen.

A production platform enforces:

  • Maximum tokens per run
  • Maximum tool calls
  • Maximum wall clock time
  • Automatic termination on overflow

Autonomy without budget control is operational risk.


2. Hard Stop Mechanism

Soft warnings are not enough.

The execution layer must enforce:

  • Max iteration count
  • Max recursion depth
  • Max retries
  • Deterministic termination conditions

The stop condition cannot depend on the model deciding to stop.

The platform decides when the agent stops.


3. Tool Gateway Layer

Agents must never call external systems directly.

Instead, every external interaction flows through a Tool Gateway:

Agent → Tool Gateway → External System

The gateway is responsible for:

  • Parameter validation
  • Schema enforcement
  • Rate limiting
  • Policy evaluation
  • Circuit breaking
  • Audit logging

This is the trust boundary.

Without a gateway, you do not have isolation. You have hope.


4. Policy Evaluation Layer

Rules must not live inside prompts.

They must live in a policy layer.

Policy defines:

  • Which tools are allowed
  • In which context
  • With which parameter ranges
  • Under which identity
  • With which budget constraints

This can be implemented using RBAC, ABAC, or a custom rule engine. The specific technology does not matter.

What matters is separation of concerns:

The model reasons. The platform enforces.


5. Audit Log and Traceability

Every run must be reconstructable.

You need:

  • Structured logs per step
  • Tool call history
  • Input and output snapshots
  • Final state
  • Failure reason

In addition, you need correlation IDs across the entire run.

If a customer asks, "Why did the agent do this?", you must be able to answer with data, not interpretation.

Audit is not optional in enterprise environments.


6. Replay Capability

This is where most systems fail.

If you cannot replay a run, you cannot debug it.

Replay means:

  • Same inputs
  • Same policy version
  • Same tool definitions
  • Controlled environment

You may not get bit level determinism because the model is stochastic. But you must get architectural determinism:

  • Same sequence of steps
  • Same policy checks
  • Same termination logic

Replay enables:

  • Postmortems
  • Regression testing
  • Safe upgrades
  • Promotion of successful paths into deterministic workflows

Without replay, every incident becomes guesswork.


7. Explicit Escalation Path

Not all missions should be autonomous.

A mature platform defines:

  • When to escalate
  • How to escalate
  • What context to attach
  • Who is responsible

Escalation triggers can include:

  • Budget exhaustion
  • Repeated tool failures
  • Ambiguous decisions
  • Policy violations

Autonomy is not the goal. Controlled autonomy is.


Architectural Separation: Control Plane vs Execution Plane

A robust agent platform separates responsibilities.

Control Plane

Responsible for:

  • Mission creation
  • Budget assignment
  • Policy resolution
  • Run state management
  • Workflow versioning

This layer is deterministic and auditable.

Execution Plane

Responsible for:

  • Running the agent logic
  • Calling tools via the gateway
  • Emitting structured events
  • Respecting hard limits

The execution plane must be sandboxed and identity scoped.

The agent does not own infrastructure permissions. The platform does.


Determinism Layer: From Exploration to Routine

In early runs, the agent explores.

Once a mission path proves stable, it can be promoted into a versioned routine:

  • Fixed step sequence
  • Defined decision points
  • Structured user inputs when required

This transforms stochastic exploration into semi deterministic execution.

It also reduces cost, variance, and operational risk.

This is where an agent platform stops being experimental and starts becoming reliable.


What This Is Not

This is not:

  • A specific cloud architecture
  • A multi agent manifesto
  • A prompt engineering guide

You can implement this on any infrastructure stack.

What matters is that the contracts between layers are explicit and enforced.


A Simple Litmus Test

Before calling your system production ready, ask:

  • Does every run have a hard budget?
  • Can I kill a run deterministically?
  • Do all tool calls go through a gateway?
  • Are policies separated from prompts?
  • Can I replay any mission?
  • Do I have SLOs for success and cost?
  • Is there a clear escalation path?

If the answer to any of these is no, you are not done.


Closing Thought

Agent systems are not impressive because they can act.

They are impressive when they can act safely, predictably, and at scale.

A Minimum Viable Agent Platform is not about sophistication.

It is about discipline.

And discipline is what turns AI experiments into infrastructure.

Partager:
Retour au blog