返回博客
Executable Memory: Why the Next Step for AI Agents Is Not Better Memory, but Executable Memory
AI 发布于 February 19, 2026 作者 Sami Kalliokoski

Executable Memory: Why the Next Step for AI Agents Is Not Better Memory, but Executable Memory

Most discussions around AI agents focus heavily on memory.
Vector databases, RAG, and conversation history are often treated as the core foundation of intelligent systems.

However, in real production environments, this is not enough.

An agent may remember what was said, but it does not remember what actually worked.
It re-solves the same task every single time:

  • with a different reasoning path
  • different tool chains
  • different step ordering
  • slightly different outputs

This leads to higher latency, higher costs, and unpredictable behavior.

To address this gap, I use an approach I call Executable Memory.


The Core Problem: Agents Remember Knowledge, Not Execution

Modern AI agents typically store:

  • conversation history
  • retrieved documents (RAG)
  • summarized past context

This helps answer:

"What do I know?"

But production systems care more about:

"How should this task be executed reliably?"

An agent can know everything about a project, yet still choose a different execution path every time it performs the same operational task. This variability becomes a serious issue in real-world systems that require consistency, auditability, and predictable outputs.


What Is Executable Memory?

Executable Memory is a memory layer that stores successful execution paths instead of just textual context.

Instead of saving only the final answer, the agent stores:

  • execution steps
  • tool calls
  • parameters
  • decision points
  • dataflow between steps
  • validated outcomes

These are then converted into reusable routines that can be executed deterministically.

Memory shifts from passive context to an active execution layer.


From Prompts to Routines: A Natural Evolution

In practice, most agent systems evolve through similar stages:

  1. Free-form prompting
  2. Tool calling
  3. Workflows
  4. DSL or scripted routines
  5. Self-learning routines (Executable Memory)

Prompt-based systems work well for exploration, but they do not scale efficiently for:

  • repetitive tasks
  • multi-system integrations
  • business-critical operations

This is where routines and executable memory become essential.


Real-World Example: Wendy and Construction Site Status Snapshots

Consider an AI agent (Wendy) operating in a construction environment with the following mission:

"Fetch 360 site images, compare them with the project schedule, and generate a current progress snapshot."

The data typically comes from multiple systems:

  • 360 camera platforms
  • project scheduling systems
  • possibly BIM and project management tools

First Execution Without Executable Memory (Dynamic Reasoning)

When the agent handles this mission for the first time, it performs dynamic planning:

  1. Identify project context
  2. Fetch 360 images for the selected date range
  3. Retrieve the project schedule
  4. Align image timestamps with planned tasks
  5. Analyze visual progress
  6. Detect deviations from the plan
  7. Generate a status summary

This requires multiple LLM reasoning loops, tool calls, and large contextual processing on every run.
It works, but it is expensive, slow, and non-deterministic.


Learning From Success: Routine Creation

After the same mission is executed successfully multiple times, the agent can detect a stable execution pattern and convert it into a reusable routine.

Example routine (DSL-style):

routine: create_site_status_snapshot
inputs:
  project_id: string

steps:

  - ask_user:
      question: "Which date range should be used for the snapshot?"
      var: date_range

  - fetch_360_images:
      project_id: ${inputs.project_id}
      date_range: ${date_range}
      output: images

  - fetch_project_schedule:
      project_id: ${inputs.project_id}
      output: schedule

  - align_images_with_schedule:
      images: ${images}
      schedule: ${schedule}
      output: aligned_progress_data

  - detect_progress_vs_plan:
      aligned_data: ${aligned_progress_data}
      output: progress_analysis

  - generate_status_summary:
      analysis: ${progress_analysis}
      date_range: ${date_range}
      output: status_report

This is not just a memory entry.
It is an executable, deterministic workflow stored in the agent's memory.


Explicit Dataflow: How One Step Drives the Next

A key advantage of DSL-based routines is explicit dataflow.

Each step produces structured outputs that directly influence subsequent steps:

  • images feed the alignment step
  • schedule defines the comparison baseline
  • aligned_progress_data drives analysis
  • progress_analysis affects later decisions

Example with a data-driven decision point:

- decision:
    condition: ${progress_analysis.delay_percentage} > 10
    if_true:
      - ask_user:
          question: "Significant delay detected. Should we run a deeper area-specific analysis?"
          var: area_filter
    if_false:
      - set_var:
          area_filter: null

Here, the output of one step deterministically shapes the next execution path without requiring the LLM to redesign the logic each time.


Decision Points and Human-in-the-Loop Execution

In real operational environments, full autonomy is not always desirable.

DSL routines can include explicit decision points where the agent:

  • asks for constraints
  • requests clarifications
  • confirms critical parameters

For example:

  • date range selection
  • floor or zone filtering
  • analysis depth

This creates controlled autonomy instead of blind automation, while keeping the workflow auditable and structured.


Skills vs Executable Memory: Different Layers, Not Competitors

Many agent frameworks rely on predefined skills such as:

  • fetch data
  • analyze images
  • generate reports

A skill answers:

"What can the agent do?"

Executable Memory answers:

"How should this task be executed optimally?"

Key distinction:

  • Skill = atomic capability
  • Routine = learned orchestration

An ideal architecture is layered:

  • Tools (APIs, integrations)
  • Skills (logical capabilities)
  • Routines / Executable Memory (optimized workflows)
  • LLM (planner and interpreter)

Why DSL + Python UDF Instead of Pure Python

A critical architectural choice is how routines are represented.

Allowing an LLM to generate raw Python for every task maximizes flexibility, but introduces:

  • larger error surface
  • weaker auditability
  • higher security risks
  • inconsistent execution structures

A DSL + Python UDF model separates responsibilities:

  • DSL handles orchestration, control flow, and dataflow
  • Python UDF handles heavy computation and specialized logic

For example:

  • DSL defines the workflow and decision logic
  • Python UDF performs image analysis, scoring, and advanced processing

This hybrid model preserves flexibility while maintaining structure and safety.


Validation, Security, and Auditability

DSL-based execution enables:

  • schema validation before runtime
  • restricted tool access
  • explicit inputs and outputs
  • step-level logging and traceability

In contrast, fully LLM-generated code is harder to:

  • validate upfront
  • audit systematically
  • sandbox safely

This distinction becomes critical in enterprise and multi-integration environments.


Reduced Hallucinations and Increased Determinism

When an agent moves from free-form reasoning to routine execution:

  • planner loops are reduced
  • tool chains become stable
  • variability decreases
  • hallucination surface area shrinks

In production systems interacting with multiple data sources, determinism is often more valuable than maximum flexibility.


Self-Learning Routines as Procedural Memory

Executable Memory effectively acts as procedural memory for AI agents.

Instead of only remembering information, the agent remembers how to perform tasks.

A routine can be automatically proposed when:

  • the task is frequent
  • success rate is high
  • the tool chain is stable
  • outputs are consistent

Over time, the agent evolves from reactive reasoning to optimized, repeatable execution.


Conclusion: From Knowledge Memory to Executable Memory

Traditional agent memory answers:

"What do I know?"

Executable Memory answers:

"How should this be executed reliably, repeatedly, and efficiently?"

When an agent like Wendy converts successful missions into routines:

  • latency decreases
  • costs drop
  • stability improves
  • auditability increases
  • systems scale better in production

The next major step in AI agents is not just better memory.
It is memory that can be executed.

分享:
返回博客