AI Agent Memory and Checkpoints
Actus · October 1, 2026
AI Agent Memory and Checkpoints
An AI agent becomes operationally useful when it can continue work instead of starting from zero. Memory preserves durable business context. Checkpoints preserve the exact progress of a multi-step or recurring workflow. Together, they let an agent resume research, avoid duplicate outreach, and carry decisions from one run to the next.
This guide explains the difference, shows how to design both, and covers the failure modes that make agent workflows unreliable.
Why Stateless Runs Break
A stateless assistant answers one prompt and forgets the operation afterward. That is acceptable for a one-off question. It is a poor fit for business workflows.
Consider a weekly lead workflow. The agent finds 100 companies, qualifies 40, drafts outreach for 25, and then the run stops. Without saved progress, the next run may:
- Research the same companies again
- Contact someone who already replied
- Lose the qualification evidence
- Repeat rejected prospects
- Miss the remaining 60 records
- Produce a report that cannot be reconciled
The problem is not only inconvenience. Duplicate outreach damages trust, and missing checkpoints make it impossible to know what was actually completed.
Memory and Checkpoints Are Different
Memory
Memory stores facts that should influence future work. Examples include:
- The company serves local service businesses
- The preferred tone is direct and practical
- The ideal customer is an owner-operated contractor
- A prospect asked not to be contacted
- The approved sending address is a specific mailbox
- A workflow should run only after a defined rest period
Memory answers, "What should the agent know next time?"
Checkpoints
A checkpoint stores structured progress for a specific workflow. Examples include:
- Records processed
- Records remaining
- IDs already contacted
- Drafts awaiting approval
- The last successful record
- Errors and retry counts
- Artifacts produced for the next stage
A checkpoint answers, "Where did this job stop, and what remains?"
Using memory for batch progress makes retrieval vague. Using checkpoints for permanent business facts makes them hard to reuse. Keep the two systems separate.
What Belongs in Durable Memory
Store facts that are stable and useful across workflows:
- Offerings and pricing boundaries
- Buyer personas and disqualifiers
- Brand voice and words to avoid
- Approval rules
- Important links and booking routes
- Team responsibilities
- Known compliance constraints
- Preferences confirmed by the user
Write facts specifically. "User likes short emails" is less useful than "Initial outreach should be under 120 words, reference one verified observation, and avoid unsupported performance claims."
Do not store secrets in article content, reports, or broadly shared memory. Credentials and tokens belong in the proper connected-account or secret-management mechanism.
What Belongs in a Checkpoint
A useful checkpoint is small, structured, and easy to resume.
For a lead workflow:
- Workflow name
- Run date
- Source list identifier
- Targeting criteria
- Processed company IDs
- Qualified company IDs
- Contacted person IDs
- Bounced or opted-out addresses
- Drafts pending approval
- Last completed stage
- Next record to process
- Error summary
For publishing:
- Published titles
- Slugs or post IDs
- Primary keywords
- Publication dates
- Cover image URLs
- Failed items and error status
The next run should be able to continue using only the checkpoint plus the current source data.
Design for Idempotency
Idempotency means repeating a step does not create unwanted duplicates.
Every record needs a stable identifier. A company website, email address, CRM ID, or post slug is better than a name alone. Before creating or sending anything, the workflow checks whether that identifier already appears in the checkpoint or destination system.
Examples:
- Before outreach, check person and company against contacted IDs
- Before creating a CRM record, search by domain and email
- Before publishing, compare the title and slug with confirmed posts
- Before retrying a failed request, confirm it did not succeed on the first attempt
A retry should resume the failed item, not repeat every successful item.
A Resume-Safe Workflow Pattern
Use stages with explicit done conditions.
Stage 1: Source
Collect the candidate batch and assign stable IDs. Done when the source list is saved.
Stage 2: Qualify
Process each candidate and save the decision plus evidence. Done when every candidate is marked qualified, rejected, or needs review.
Stage 3: Prepare
Create drafts only for approved qualified records. Done when each draft is linked to its record.
Stage 4: Execute
Send, publish, or update the external system only after approval. Save the destination ID immediately after each confirmed success.
Stage 5: Report
Summarize confirmed results, failures, and remaining work. Done when the report matches checkpoint counts.
After each record, update the checkpoint. Do not wait until the end of a long batch. A run can stop at any time, and the latest completed record should still be recoverable.
Handling Partial Failure
External systems fail temporarily. A sound workflow records:
- The item identifier
- The action attempted
- The error and status code
- Whether the action might have succeeded despite the error
- The next retry time
- The number of attempts
For rate limits, respect the stated retry delay. For uncertain network failures, verify the destination before retrying. For a publishing API, check whether the post exists before submitting it again. For email, do not resend until the delivery state is known.
Three states are useful:
- Confirmed success
- Confirmed failure
- Unknown, requiring verification
Unknown is not the same as failed.
Avoiding Duplicate Outreach
Deduplicate at several levels:
- Same company with multiple locations
- Same domain with different name spellings
- Same person using multiple email addresses
- Generic inboxes shared by several contacts
- Records already present in the CRM
- People who opted out or replied negatively
Store normalized values. Lowercase domains and email addresses, remove tracking parameters from URLs, and keep both the original and normalized forms when useful.
When a person replies, the checkpoint or CRM state must stop future sequence steps. A reply is a state change, not just another event.
Context That Improves Quality
Memory should make outputs more relevant without becoming a dumping ground.
Good context includes:
- "Disqualify companies that only serve commercial new construction."
- "Use the booking link in follow-ups after a positive reply."
- "Refer to the business as a service company, not an enterprise platform."
- "Do not mention pricing unless the user asks."
Poor context includes:
- Entire transcripts with no summary
- Temporary batch notes
- Contradictory instructions
- Sensitive personal data unrelated to the workflow
Review memory periodically. Remove outdated offers, old priorities, and conflicting rules.
Multi-Stage Handoffs
Checkpoints can pass work between stages or specialized agents.
Example:
- Research agent saves qualified leads and evidence
- Audit agent adds website findings
- Offer agent creates the angle
- Messaging agent drafts outreach
- Sending stage records confirmed delivery
Each stage should receive only the fields it needs and return a structured result. The handoff includes record IDs and source evidence so the next stage can verify the work.
This prevents one long workflow from hiding errors. If qualification is weak, fix that stage before paying to generate messages.
Observability and Auditing
A checkpoint should support a human-readable report:
- Started with 100 records
- 62 processed
- 28 qualified
- 20 drafts approved
- 18 confirmed sent
- 2 failed because of invalid addresses
- 38 remain for the next run
The counts must reconcile. If the report says 18 sent, the checkpoint should contain 18 confirmed destination IDs.
Also store enough evidence to audit a decision. For qualification, save the URL and observation. For publishing, save the post ID and slug. For outreach, save the recipient, message version, and confirmation state.
Retention and Privacy
Not all progress data should live forever.
Define retention by workflow:
- Active batch state: keep until complete
- Suppression and opt-out lists: keep as long as required to honor them
- Routine logs: keep for a limited operational period
- Sensitive documents: store only when needed and restrict access
Avoid placing personal data in titles, public reports, or content generated from checkpoints.
Testing the Resume Path
Do not assume a workflow can resume because the logic sounds correct. Test it.
- Process a small batch
- Stop after partial completion
- Restart the workflow
- Confirm completed items are skipped
- Confirm the next item is processed once
- Introduce a temporary failure
- Verify the retry does not duplicate success
This test reveals whether identifiers, destination checks, and checkpoint updates actually work.
Common Design Mistakes
Saving Only a Summary
"Processed about half the leads" cannot resume anything. Save record IDs.
Updating the Checkpoint Too Late
If the run fails before the final save, all progress is lost. Update after each confirmed record.
Confusing Attempted with Completed
A request that times out is not confirmed. Verify it before recording success.
Keeping Contradictory Memories
Old and new instructions can produce inconsistent behavior. Replace outdated facts explicitly.
Mixing Workflows in One Key
Separate checkpoints for lead research, publishing, and follow-up. A shared blob becomes difficult to repair.
Getting Started with Actus Agent
Actus Agent supports durable memory for business context and structured checkpoints for resumable workflows. Use memory for stable facts such as audience, voice, offers, and approval rules. Use a checkpoint for batch progress, processed identifiers, confirmed outputs, and remaining work.
Start with one recurring workflow. Define its stages, stable record IDs, and done conditions. Save progress after every confirmed item. On the next run, read the checkpoint first and process only what remains.
Memory makes an agent more relevant. Checkpoints make it reliable. A workflow needs both to move from a promising demo to an operating system you can trust.