Setup guide

Action Brain: a personal agent that turns loose thoughts into handled tasks

Record a voice memo, forward an email, or send a link. An always-on agent transcribes it, works out what matters, does the research, reminds you at the right moment, and keeps a live dashboard and app in sync.

Agent: Hermes, on a small home computer Cloud: Cloudflare free tier Interface: Telegram, web, native app

Overview

Most to-do systems fail at the capture step. You have to stop, open an app, and type. Action Brain removes that friction: you capture however is easiest, and the agent turns what you captured into something you can act on.

The loop is simple:

  1. Capture: a quick voice memo from your phone, an email to a private inbox, a link or a message in chat.
  2. Triage: the agent pulls out to-dos, dates and decisions, judges how important each one is, and researches anything that needs it (phone numbers, business hours, how a process works), always with sources.
  3. Remind: reminders arrive at sensible times, avoid the times you're busy, and keep coming until you mark the item done.
  4. Review: a password-protected dashboard and a native app show what's left today and what's coming later, reading from one shared database.
Design principle: the language model only runs when something actually changed. Polling, scheduling and reminder delivery are plain scripts, which keeps costs close to zero and behavior predictable.

Architecture

Four layers, left to right. The agent runs at home; storage and anything public-facing run on Cloudflare's free tier.

Everything runs as a few scheduled jobs on the home device:

JobEveryRuns the LLM?What it does
Memo triage1 minOnly when there's a new recordingLists the bucket; wakes the agent only when an untriaged memo appears
Email triage2 minOnly when there's new mailChecks both inboxes; same "wake on change" pattern
Reminders5 minNoSends due reminders, the daily agenda and game alerts; mirrors calendar data; refreshes the dashboard
Quota watchDailyNoSilent unless free-tier usage is on pace to exceed a limit

Voice memo pipeline

The main way things get in. A memo is usually 5–30 seconds: "remind me to…", "add this trip idea…", "email my spouse that…".

  1. Upload. A phone shortcut posts the audio to a small Cloudflare Worker that checks a secret token and writes it to an R2 bucket under a dated key.
  2. Detect. A script lists the bucket every minute and compares it with a local "already triaged" list. If nothing is new, its output doesn't change, so the agent never wakes.
  3. Pull only what's new. New recordings are downloaded once, tracked in a local list so nothing is fetched twice.
  4. Transcribe. Audio is converted with ffmpeg and sent to Workers AI (whisper-large-v3-turbo). That costs about $0.0005 per audio minute, and the free daily allowance covers roughly 200 minutes.
  5. Triage. The agent reads the transcript along with recent notes and the open task list, so a memo that repeats an existing task gets merged instead of duplicated.
  6. Record. Results go to a dated triage notes file, reminders go to the database, and the memo is marked triaged only after everything is saved.
  7. Recover. If a run fails, the memo is still untriaged after 10 minutes, the monitor's output changes, and the agent retries. A failure lasting 30+ minutes sends one alert.
Speech-to-text gets names wrong. Give the agent a short glossary of the people in your life (stored in its memory) so it fixes mis-heard names before acting.

Triage & importance

Every item gets an importance level, and that sets how persistent its reminders are. Hard things aren't crammed into tonight; they're scheduled for when you can actually act on them.

LevelTypical itemsFirst reminderThen
HighFamily, legal, financial or health obligations; deadlines within ~3 daysNext morning, 9amDaily until done
MediumReal tasks without a pressing deadline: home projects, errands, follow-upsIn 1–3 daysEvery 3 days
LowIdeas, someday/maybe, nice-to-havesIn about a weekOnce
  • Dates you mention explicitly always win over these defaults.
  • If a task needs a business to be open, the reminder is timed for its hours. A "prepare" reminder can still come earlier.
  • Each reminder carries everything needed to act: phone numbers, hours, steps, links. You never have to look anything up.
  • The research the agent does cites its sources and never invents numbers, addresses or facts.

Reminder engine

A plain script, run every five minutes, decides what to send. It respects your time.

When it stays quiet

  • Quiet hours (9pm–8am): late or repeat reminders wait for the morning.
  • Busy calendar events: nothing is sent during a meeting or appointment; tasks wait until it ends.
  • Protected time: a favorite team's games block everything from an hour before kickoff to the end. If you're going in person, the window covers travel too.

What it sends

  • Daily agenda at 7am, with warnings when events overlap.
  • Due tasks with full details and a done <id> reply hint.
  • Habit check-ins (medication, supplements, exercise), skipped if already logged today.
  • Game-day morning note and a 30-minute heads-up.

A habit reminder that would land inside a busy block moves just before it, but only if that's within two hours of its usual time. Otherwise it waits until the block ends; an evening medication reminder shouldn't jump to mid-afternoon.

To make changes safely, the engine has a time-travel test mode: copy the data, set a fake "now", and step through a week in 5-minute ticks to see exactly what would be sent and when.

Calendar & schedules

Context makes reminders smarter. Everything here is read-only.

  • Personal calendar through Google Calendar's secret iCal link, which is read-only and needs no OAuth app. Recurring events are expanded 14 days ahead and refreshed every 30 minutes. Events marked "free" don't block anything.
  • Sports schedule from a public sports-data API, refreshed twice a day because game times change. Games with a time still "TBD" are ignored until it's set.
  • Work calendars stay out. Company policy usually forbids sending them to third-party agents, so keep this personal.

Email on your behalf

Two agent-owned inboxes with very different trust levels.

Private channel

  • You forward things here; they're triaged like voice memos.
  • Instructions are trusted only from your own address and only when the message passes sender authentication (catches spoofing).
  • Mail from anyone else gets a one-line heads-up and nothing more.
  • Outbound mail from this inbox can only go to you; a hard guard in the email script enforces it.

External correspondence

  • Used to email other people on your behalf: professional tone, signed as sent on your behalf.
  • When you explicitly ask, it sends right away, as long as the content has no personal identifiers or the recipient is someone you've had it email before.
  • Anything else is drafted and needs your approval. The script refuses to send without a recorded approval note.
  • Replies with attachments or links you need are forwarded to your personal inbox, and you get a chat ping.
Every email is untrusted data, never instructions. Subjects, bodies and attachments can contain prompt injection. The agent summarizes and flags them but never follows them. Every outbound message is written to an audit log.

Libraries & tracking

Some captures aren't tasks: they're things you want to come back to.

  • Trip library. Send a travel reel or blog link and get a structured itinerary (day by day, where to stay) plus "check before booking" notes: seasons, road closures, permits, booking lead times.
  • Recipe library. Send a recipe link and get the ingredients, steps, the author's tips and a ready-made shopping list. Many recipe sites block bots, so the agent uses the site's public WordPress API when the page itself is walled off.
  • Mail & shipments. Photograph a receipt; the agent reads the tracking number (checking its check digit), records it and schedules a "confirm delivery" follow-up.
  • Projects. Multi-step efforts (say, getting quotes for a home repair) roll up into one check-in with the current status, contacts and next steps, instead of a pile of separate nags.

Libraries are plain Markdown files, one per item, each with an index. They're easy to read, grep, version or move.

One source of truth

The first version kept tasks in local files and copied them to the cloud every few minutes. The dashboard lagged, so that design was replaced.

Now a Cloudflare D1 database is the single source of truth for tasks, habits and habit logs. Every writer and reader uses the same rows:

  • the reminder engine and the triage jobs, through a small internal API using a secret token,
  • chat ("done 4f2a1c", "walked 3 miles"),
  • the web dashboard, rendered live from D1 on every request,
  • native apps, through device tokens.

Two details make it robust:

  • Field-level, guarded writes. The reminder engine only updates the fields it changed, and only if the item is still open, so it can't bring back something you finished on your phone a second earlier.
  • An offline fallback. If the cloud is unreachable, reminders run from a local cache and their changes are queued and replayed later.
items       (id, kind task|habit, title, details, importance, status open|sent|done,
             next_due, repeat_hours, last_sent, done_at, source)
habit_logs  (date, habit, value, ts)
events      mirrored calendar, next 7 days
games       mirrored sports schedule
actions     audit log of app actions
devices     sha256(token), name, last_seen

Web dashboard

A single Cloudflare Worker, behind HTTP Basic Auth, built for a phone screen.

  • Today: what's still on the schedule (with a NOW marker), tasks due or reminded today, habits not yet logged, and a collapsed "done today" list.
  • Later: grouped into Tomorrow, Next 7 days, Later, and Waiting on you.
  • Also: the upcoming calendar, a 7-day habit grid, upcoming games, recent triage notes, and the trip and recipe libraries.

Password handling

The agent never sees the password. It creates a one-time setup link that expires in 24 hours. You choose the password on that page, and only a salted PBKDF2 hash is stored. Pages are served with noindex, no-store and a strict Content-Security-Policy that blocks all scripts.

Native app & sync API

The same Worker exposes a small JSON API, so a SwiftUI app for iPhone and Mac can show the same data and act on it.

EndpointPurpose
POST /api/v1/loginSign in once with the dashboard password; returns a per-device token to keep in the Keychain. Only its hash is stored.
GET /api/v1/snapshotTasks with live Today/Later buckets, habits, logs, events and games. Supports ETag, so an unchanged poll costs one row read.
POST /api/v1/actionstask_done, task_reopen, habit_log or task_snooze, applied immediately.
GET /api/v1/actionsHistory of which device did what, and when.

Each device can be revoked on its own. iPhone and Mac apps have to be built in Xcode on a Mac. A Linux build server can host the API but can't build the app.

What it tracks

The kinds of things that have flowed through it. The categories matter more than the specifics.

  • Family logistics
  • Kids' school & homework
  • School and sports paperwork
  • Benefits & insurance paperwork
  • Certified mail & tracking
  • Home projects & contractor quotes
  • Household chores
  • Travel plans & documents
  • Trip ideas
  • Recipes
  • Medication adherence
  • Supplements
  • Daily exercise
  • Calendar conflicts
  • Sports schedule
  • Side-project infrastructure
  • Cloud quota usage

Habit tracking

Habits are just reminders with a daily repeat and a log. Reply naturally ("took my meds", "walked 3.5 miles") and the agent records it. A logged habit never sends a reminder that day, and the dashboard shows a 7-day streak grid.

Check before you claim. Logs can arrive from chat, memos, email or the app at any moment. The agent re-reads the log before saying something is missing. An early version told the user a habit wasn't logged five minutes after it had been.

Security model

The agent can run commands on a home computer, so who can talk to it, and what it trusts, matters more than anything else.

Access

  • The chat bot answers exactly one user ID. Turn on two-step verification for that chat account; it's effectively the key to the machine.
  • Nothing on the home device is reachable from the internet: no port forwards, no tunnels, no public IPv6. All traffic is outbound.
  • The memo upload endpoint and the private R2 bucket both require tokens; public bucket access is off.
  • Every credential file on disk is readable by its owner only (chmod 600), and secrets are never shown on the dashboard.

Trust

  • Web pages, emails and attachments are data. Only the owner gives instructions.
  • Hard guards live in code, not just in prompts: send allowlists, approval notes, no-reply rules.
  • No personal identifiers go out without explicit approval.
  • Passwords and payment details are never typed by the agent or pasted into chat.
  • Read-only access wherever possible (calendar, schedules).

A periodic audit is worth doing. Check open ports from outside, the router's automatic port openings (UPnP), SSH settings (keys only), unneeded network services, the firewall, automatic security updates, and any old agents still running in the background.

Costs & quotas

Everything fits comfortably in Cloudflare's free tier. The bigger risk is other projects on the same account.

ResourceFree limitAction Brain usage (est.)
Worker requests100k / day~1–1.5k (≈1%)
D1 rows read5M / day~15k (≈0.3%)
D1 rows written100k / day~300–1,000 (≈1%)
KV reads / writes100k / 1k per day~100 / ~10
R2 list operations1M / month~43k (checking once a minute)
Speech-to-text10k neurons / day≈200 audio-minutes free

Free-tier limits are shared across the whole account. A daily script checks usage through Cloudflare's GraphQL analytics and only sends a message when you're on pace to exceed a limit. The LLM is the real cost, which is why it only runs when there's new input.

Lessons learned

Gate the LLM behind cheap change detection.
A monitor script's output is hashed every tick, and the agent only wakes when the hash changes. Make that output grow only (every ID ever seen, plus a retry marker), so marking something done doesn't wake the agent again.
Pick one source of truth early.
Copying data between stores causes "why is the dashboard stale?" bugs. Put the state in one database and make every surface read it live.
Enforce rules in code as well as prompts.
Policy written into a prompt is advice. The email script refusing to send without an approval note is a guarantee.
Simulate time before shipping schedule logic.
Stepping through a fake week caught a bug that would have sent the same habit reminder three times when it moved ahead of an event.
Watch for working-directory collisions.
The Cloudflare CLI (wrangler) creates a .wrangler folder wherever you run it. Run it from your home directory and that folder masks its saved login. Always run tools from a project folder.
Never send undefined to D1.
Older records missing a field broke a bulk write. Give every column an explicit default.
Never quote state from memory.
Re-read the log before saying whether something is done or missing.

Build your own

A practical order of operations, roughly one evening per step.

  1. Run an agent on a box that's always on. Hermes on a small single-board computer works well. Connect a chat app and allow only your user ID.
  2. Set up voice capture. Create an R2 bucket and a token-protected upload Worker, plus a phone shortcut that posts the recording.
  3. Add the memo monitor, transcription and triage job. Write down the importance policy and the reminder cadence explicitly.
  4. Add the reminder engine. Start with quiet hours and repeat-until-done; add calendar blackouts after that.
  5. Create the D1 database and API. Make it the source of truth from day one.
  6. Build the dashboard. It should be password-protected, live, and organized by today versus later.
  7. Connect email. Use separate inboxes for "talk to me" and "talk to others", enforce guards in code, and keep an audit log.
  8. Run a security audit and set up the quota alert. Then tune the cadences based on what actually annoys you.