GUIDES / PI DURABLE

Build a durable Pi agent that can recover after a crash

Use normal Pi for an interactive coding task you can supervise and restart. Choose experimental Pi Durable when your application must recover unfinished work from persistent state. Cloudflare PiHarness is an optional beta runtime that adds Durable Object storage and wake-up handling. Before enabling recovery, decide which tools may replay and how to reconcile external writes: recovering an agent does not guarantee that every side effect happens exactly once.

Choose the runtime and recovery boundary ↓

1. Start with the simplest runtime that fits

AgentSkillsHub editorial selection guidance: normal Pi usually fits a short local coding task when a developer is present, external effects are limited and restarting is inexpensive. It can also be scripted; “normal” does not mean interactive-only. A saved coding session is useful, but it is not the same contract as automatically recovering a pending tool task.

Consider Pi Durable for hours-long coding or research work, asynchronous browser/network dependencies, a human reply arriving later, or a machine restart between steps. Scheduled continuation needs a host that actually wakes the work; a storage file does not schedule itself. If an existing job system already owns recovery, avoid introducing a second owner without a concrete need.

Do not choose Durable yet if you cannot identify pending actions, classify safe retries or reconcile an uncertain external write. Durability is not a permissions layer. High-impact account, financial or production actions still need approval and verification.

The choices are a coding-agent system, an experimental durable library, and an optional Cloudflare runtime integration. They are not interchangeable product editions.

2. Keep the version and package layers separate

Current published Pi version: 1.0.2, checked 2026-10-04. Pi 1.0.0 on October 1 is the major milestone; 1.0.1 followed October 3 and 1.0.2 October 4. The latest patch adds sampling overrides by thinking level, not a new recovery guarantee. Current changelog · 1.0 milestone.

Pi describes itself as a minimal, extensible agent harness. The repository separates these responsibilities:

  • Model abstraction — pi-ai: model/provider access used by higher layers.
  • Agent core — pi-agent-core: agent-loop and state machinery.
  • Coding-agent CLI — pi-coding-agent: the terminal application and its SDK surfaces.
  • Extensions: executable customizations that add behavior, tools and hooks.
  • Skills: task instructions/resources loaded when relevant, distinct from executable extensions.
  • Codemode: JavaScript orchestration of available tools. Its script sandbox does not confine every host-side tool.
  • MCP: external tool/resource connections supplied through the coding agent’s built-in extension.
  • pi-durable: a separate harness for persisted conversations, application documents and checkpointed tasks, built on model access and Chord state.

The repository and Durable package use the MIT license. That does not make model calls, hosting or third-party MCP services free. License · Durable package.

3. Configure MCP, then choose tool exposure

Pi 1.0 supports MCP through a built-in coding-agent extension; the changelog already listed this in 0.99.0. Do not describe MCP as a capability of every Pi layer or assume a Durable host automatically imports the terminal’s extensions and config. Release history · MCP documentation.

The documented transports are stdio and streamable HTTP, not legacy SSE. User configuration lives in ~/.pi/agent/mcp.json; trusted projects may use .pi/mcp.json. Full same-name project entries replace user entries. A partial entry without command, URL or type overrides only enabled/exposure/toolExposure settings.

Use pi mcp add with the selected server’s actual command or URL; pi mcp list and interactive /mcp show connection state. Reload after an out-of-session change. OAuth is supported through discovery or explicit settings; pi mcp login and logout manage sign-in. Pi 1.0 hardens issuer/scope handling, and 1.0.1 adds CIMD registration. Check the server’s required flow rather than copying another client’s credentials.

Tool exposure controls context, not trust. The four MCP modes below govern discovery/declaration. No token-savings percentage is claimed here. A tool connection does not prove the server is safe or that a task completed. Exposure and permissions.

codemode
Default: discover and invoke from scripts; tools are not all declared to the model.
deferred
tool_search loads matching declarations for the next model call.
direct
Declared to the model and also callable from codemode.
hidden
Registered but unavailable to the model and scripts.

4. Normal Pi vs Pi Durable

“Normal Pi” below means the coding-agent CLI/SDK without adding the Durable task runtime. It has saved sessions and resume controls; the distinction is recovery of in-flight execution, not whether any history can be stored. The comparison combines the Pi repository and Durable contract; suitability judgments are AgentSkillsHub editorial guidance.

Pi Durable is EXPERIMENTAL. Its API may change without notice. Pin the implementation you evaluate and verify compatibility before reopening persisted work after an upgrade.

Normal Pi vs Pi Durable — execution and recovery
DimensionNormal PiPi Durable
Process lifetimeNormal PiRuns in the CLI/SDK host; an exit stops that process.Pi DurableA host can reopen persisted work after its process dies.
StateNormal PiSaved coding sessions and configurable extensions.Pi DurableCommitted transcript entries, JSON documents, inbox and durable tasks.
Crash recoveryNormal PiResume session context; do not assume automatic recovery of an in-flight tool.Pi DurableRecover pending tasks from checkpoints in the same storage.
Restart behaviorNormal PiRestart the application and choose the saved session.Pi DurableReopen storage, reinstall definitions and resume the scheduler.
Human steeringNormal PiInteractive steering/follow-up controls in the coding workflow.Pi DurablePersisted queued input; steer after the tool round, follow up after the answer.
Tool requestsNormal PiExecute in the current process; application-specific retry handling.Pi DurableIntent committed before execution; committed output/result records.
CheckpointingNormal PiSession history is not a durable task-state contract.Pi DurableTask phases persist checkpoints; not a process-memory snapshot.
StorageNormal PiCoding-agent session storage and host-owned files.Pi DurableMemory is ephemeral; SQLite/JSONL provide persistent backends with limits.
Retry behaviorNormal PiInspect what already happened before a manual or host retry.Pi DurableInterrupted models may repeat; interrupted tools replay only if declared safe.
Best fitNormal PiSupervised, short or easily repeated local coding work.Pi DurableLong waits, recoverable multi-step work and host restarts.
MaturityNormal PiPi 1.0 milestone; reviewed current patch 1.0.2.Pi DurableEXPERIMENTAL: APIs can change between releases.

5. Design persistent state before designing retries

Pi Durable stores immutable transcript entries alongside typed JSON documents. Documents track items such as agent choices, inbox, live work and usage. An atomic storage commit can change entries, documents and tasks together. A task is a durable state machine with checkpoints; it is not a snapshot of every JavaScript stack frame. Concepts and storage.

Choose the backend deliberately: MemoryStorage is ephemeral. Node SQLite uses WAL with synchronous NORMAL: process-crash recovery does not guarantee the newest writes survive power or host failure. JSONL offers an fsync option before commit markers. Current storage has one process owner and no cross-process locking. Storage limits.

For recovery, reopen the same persistent storage, install the required extension definitions and resume the scheduler. Recreating a blank store or changing task definitions is not recovery of the original task. Preserve task/conversation identity and inspect the expected state before continuing.

Treat the persisted transcript and state as the recovery record. A UI event stream is not a complete journal: a slow or reconnecting observer may receive a fresh snapshot instead of every historical update. Export an audit trail separately if your application requires one. Watching a conversation.

6. Separate admission, model retry and tool replay

The current Durable README documents three different boundaries:

  1. Submission admission: retrying a submission with the same requestId returns the existing submission. Keep that identity stable across your own retries.
  2. Model recovery: an interrupted model request may run again. Stored partial output does not make the request free to repeat; attempts can add usage.
  3. Tool recovery: the tool intent is committed before execution. An interrupted call reruns only when its definition declares replay: "safe". Otherwise the model receives an interrupted result with committed output so far.

Record the tool request and committed result, not just the final answer. A request marked interrupted may still have changed an external system. The model deciding a next step is not independent proof that replay is safe.

Human steering joins after the current tool round; follow-up input waits for the run’s answer. Cancelling a wait cancels that wait, not the underlying work. Build explicit stop/approval behavior at the host boundary. Tool execution · Busy conversations.

7. Side-effect boundary: recovery is not exactly-once delivery

SIDE_EFFECT_BOUNDARY — AgentSkillsHub editorial guidance. Imagine a tool triggers an external API operation successfully, then the agent crashes before recording its acknowledgement. Durable storage knows the request was intended, but the local record cannot prove whether the remote action happened. A retry can duplicate the action; refusing a retry can leave the workflow incomplete.

The documented storage/replay contract does not establish an exactly-once guarantee for payments, emails, deployments or other external mutations. Admission requestId deduplication is not a cross-system transaction. Documented boundary.

Use these editorial patterns where the destination supports them:

  • Keep a stable destination idempotency key and its scope; do not rotate it to escape uncertainty.
  • Query before retrying. Match the destination’s operation ID, receipt and current status to the intended action.
  • Read after write and verify the expected result before marking delivery complete.
  • Require explicit confirmation for a new high-impact action and manual escalation when its outcome remains unknown.
  • Declare replay safe only when the specific operation and destination contract support it. A tool named “write” is not automatically safe.

Durable execution preserves a place to resume the decision. It does not supply missing business receipts or authorize another attempt.

8. Cloudflare PiHarness is an optional beta runtime

The Agents SDK exports PiHarness from agents/harness/pi to host Pi Durable in an Agent or plain Durable Object. PiHarness is BETA; Pi Durable remains EXPERIMENTAL. Cloudflare is optional: Durable also has independent Node storage paths. Cloudflare guide · Launch note.

Pi’s transcripts, inbox and tasks live in the object’s SQLite storage. Its scheduler runs in memory. A Lifecycle job keeps an alarm while work is pending; after eviction/restart it reopens stored work and resumes. Persistence does not mean the process never dies.

The application owns client transport. After a restart, open a new event stream and send its fresh snapshot; there is no stream replay cursor. The guide also notes tool-approval limitations and that tools ignoring abort signals can delay stopping. These are runtime constraints to evaluate, not production-stability claims.

For a selected host, configure storage, the model adapter, extension registry and wake mechanism together. Do not assume a coding-agent mcp.json is automatically consumed by PiHarness. Check the host’s tool adapter and authorization explicitly. This review deployed nothing.

9. Failure example: a coding agent waits for CI

ORIGINAL EDITORIAL EXAMPLE — NOT RUN. A long-running coding agent prepares a branch, waits for CI and then continues. The goal is to review the same branch revision after a restart, not to launch another workflow blindly.

  1. Start task: persist a task ID, repository/branch, intended result and authorized action scope.
  2. Edit branch: record the exact commit revision and the work already completed.
  3. Trigger CI: after approval, save the CI run ID and a stable request/idempotency key where supported. If no receipt arrives, mark the action uncertain.
  4. Runtime restarts: reopen the same storage and required task definitions; do not resubmit the original user task under a fresh ID.
  5. Restore state: classify the pending step as read, confirmed write or uncertain write.
  6. Re-check CI: query the existing run and verify that its commit matches the saved revision. If the trigger outcome was uncertain, search/reconcile the destination first.
  7. Continue or ask: consume the verified result, or ask for renewed approval if scope, branch state or permissions changed. CI success alone is not permission to merge or deploy.

No GitHub action, CI job, deployment or model call was triggered for this example. An initial implementation test should use synthetic inputs and a fake external-action sink.

10. Recovery acceptance checklist

AgentSkillsHub editorial checklist — not a vendor certification. Test on isolated synthetic data before relying on a durable coding agent. The documented recovery semantics define the expected mechanism; the checks below define our proposed acceptance.

Before restart

  • Save task/conversation ID and the requested goal.
  • Confirm the checkpoint and expected state version are persisted.
  • Identify the pending tool call and whether its result was committed.
  • Tag external side effects and preserve any destination operation IDs.
  • Record the stable idempotency key where applicable and the current approval scope.

After restart

  • Recover the same task identity from the same storage.
  • Compare the state version and registry definitions with expectations.
  • Classify each pending action before continuing.
  • Query the external system before retrying an uncertain write.
  • Verify no duplicated irreversible effect; UNKNOWN is not evidence of none.
  • Refresh human approval after a long wait if scope or assumptions changed.
  • Confirm the final output still matches the requested goal, not just a completed task status.

If any receipt or state transition is ambiguous, preserve the checkpoint and escalate the specific uncertainty. Do not mark a recovery test passed solely because the agent produced another message.

11. Revalidate permissions when work resumes

Pi’s security documentation says the ordinary process inherits its launching user’s permissions; it does not include a general built-in gate restricting filesystem, process, network or credentials. The PiHarness guide likewise notes that Durable has no tool approval step yet. Durable state is not a sandbox or an authorization service.

AgentSkillsHub editorial guidance: trust-check MCP servers and executable extensions; constrain filesystem roots and network destinations through the host/sandbox; give credentials only the required scope. A codemode JavaScript sandbox does not isolate host-side MCP effects.

Keep raw secrets out of prompts, transcripts, logs and recovery state. Resolve credentials through the host’s approved secret mechanism. On resume, detect expiry or revocation before using a tool; do not persist a broader replacement credential to bypass a denial.

Long waits can invalidate assumptions. Re-check branch revision, destination account, tool definition, ownership and approval expiry before acting. Obtain renewed human approval for changed or high-impact scope. Our MCP server security checklist provides the related connection review.

Start with normal Pi when you only need a supervised local code task. Add durability when the recovery requirement and side-effect rules are concrete, then evaluate it against the checklist above.

FAQ

Is Pi 1.0.0 the latest version?

No. On 2026-10-04 the official changelog lists 1.0.2. Version 1.0.0 is the October 1 milestone, followed by 1.0.1 on October 3.

Does normal Pi have persistent sessions?

Yes. Saved coding sessions are distinct from Pi Durable’s checkpointed task recovery. Do not infer automatic recovery of every in-flight tool from session resume.

Where does Pi MCP support live?

The coding agent ships a built-in MCP extension. Its four documented MCP exposure modes are codemode, deferred, direct and hidden. A Durable host must select and integrate its own tools and extensions.

Does Pi Durable guarantee exactly-once external actions?

No universal guarantee is established. requestId deduplicates submission admission; an external write can succeed before its acknowledgement is stored. Reconcile the destination before retrying.

Do I need Cloudflare for Pi Durable?

No. Pi Durable has independent storage options. Cloudflare PiHarness is a beta integration using Durable Object SQLite and Lifecycle wake-up handling.

What survives a restart?

Only state persisted by the chosen backend under its durability contract. MemoryStorage does not survive process loss. Reopen the same durable storage with the required definitions; storage recovery is not a promise that every external effect can be replayed.

Can a human steer a running durable task?

Queued steering joins after the current tool round, while follow-up input waits for the answer. Cancelling a wait does not cancel the underlying work. Approval and stopping behavior still need explicit host handling.

Was this crash-recovery workflow tested?

No. This is a source-checked guide with an original failure example and editorial acceptance checklist. Pi Durable is experimental, PiHarness is beta, and no model call, CI trigger or deployment was performed.

Primary sources

Source-checked 2026-10-04. Documentation review is not a crash-recovery test. Examples and checklists here are editorial material.