Problem

A business can be run from one platform: posts go out on X, Instagram and LinkedIn, cold emails are sent, payments arrive by webhook, customers get a one-time download link by mail, and code ships through GitHub and Vercel. An agent can drive all of it. Nobody wants an agent holding those credentials directly.

Guardian is the layer in between. The agent never talks to an external service itself. It asks Guardian, and Guardian classifies the task, applies policy, asks a human when the policy says so, and only then calls the service through a connector.

running agent
     |  task request
     v
+---------------------------------------------+        +-----------+
|                  GUARDIAN                   | notify |  owner    |
|                                             |------->| Telegram  |
|   classify --> policy --+--> auto ---+      |<-------| approve / |
|                         |            |      | reply  | reject    |
|                         +--> approval|      |        +-----------+
|                              queue --+      |
|                                      v      |
|                                  dispatcher |
+--------------------------------------+------+
                                       |  connectors
    +------------+-----------+---------+--------+---------+
    v            v           v         v        v         v
   X       Instagram    LinkedIn  Zoho Mail  GitHub    Vercel
                                  (cold mail,
                                   download links)

inbound
Dodo Payments webhook ---+
                         +--> webhook receiver --> dedupe --> action
website query webhook ---+     (Redis + Postgres)    (e.g. mail a one-time link,
                                                      notify the owner)

state: Postgres (tasks, events, approvals)   Redis (fast dedupe, locks, tokens)
Where Guardian sits between the agent, the outside world and the owner.

What it does

  • Guardrail for agent actions. Every outbound task is classified and checked against policy. Low-risk work goes through, sensitive work waits for approval.
  • Approval flow. Held tasks are sent to the owner on Telegram, and the agent continues only after an answer.
  • Social publishing. Content is created and posted to X, Instagram and LinkedIn through connectors.
  • Cold mailing and delivery. Mail goes out through Zoho Mail, including one-time downloadable links sent after a purchase.
  • Payment webhooks. Dodo Payments events arrive at a receiver that deduplicates them, then trigger the follow-up, such as sending the download link.
  • Website queries. Queries submitted on the website arrive as webhooks and notify the owner.
  • Developer services. GitHub and Vercel are reached the same way, behind the same policy.

Decisions

D1Two layers of deduplication

Webhook receivers see the same event more than once. A Redis SET ... NX key drops the obvious duplicates cheaply, before anything touches the database. It is not the source of truth: keys expire and Redis restarts. The source of truth is a unique constraint in Postgres, written with INSERT ... ON CONFLICT DO NOTHING. If no row comes back, another worker owns the event.

D2Flag interrupted calls instead of retrying them

If the process dies after sending a request to a connector but before recording the answer, nobody knows whether the action happened (a mail sent, a post published). A blind retry risks doing it twice. Guardian marks the row for manual reconciliation instead, which turns an unsafe guess into a visible, bounded queue.

D3State changes are compare-and-set

Every transition is an UPDATE whose WHERE clause includes the expected current state, so two workers cannot both move the same row. A partial unique index enforces that a payment has at most one active fulfilment.

-- Simplified schema.
update fulfilment
   set state = 'dispatching', attempts = attempts + 1
 where id = $1 and state = 'pending'
returning id;

create unique index one_active_fulfilment
    on fulfilment (payment_id)
 where state in ('pending', 'dispatching');

D4Single-use tokens are claimed atomically

Reading a token and deleting it must be one step, or two requests can both read it before either deletes it. A short Lua script does both inside Redis.

Known issues

Two problems I found in review and have not fixed yet:

  • Lock release ignores ownership. The loop lock is released unconditionally. If its TTL has already expired and another worker has taken it, the release deletes someone else's lock. The fix is to store a token and compare before deleting, in one script.
  • Fixed-window rate limiting overshoots. A burst at the end of one window and the start of the next can let through close to twice the limit. A sliding window or token bucket removes the edge.

Some adapters are only partly wired behind their interfaces. That is next after the two fixes above.

Longer write-up on the idempotency design →

← All projects