Skip to main content

Command Palette

Search for a command to run...

I Built an AI Agent That Reads My Email and Replies as Me โ€” Email It and See

Updated
โ€ข9 min readโ€ขView as Markdown
I Built an AI Agent That Reads My Email and Replies as Me โ€” Email It and See

There's a live email address, right now, sitting behind a real AI agent I built:

๐Ÿ“ง askrajat.ai@gmail.com

Email it a question about my background, my hiring rates, my technical stack, or whether I'm available for freelance work โ€” and a deployed pipeline running on Google Cloud will read it, reason about it, retrieve grounded facts about me, draft a reply, independently check that reply is safe to send, and email you back. No human in the loop. If you're reading this and wondering whether I can actually build production AI systems โ€” this is the demo, not a screenshot of one.

This post is about how it works, the decisions behind it, and โ€” more usefully โ€” the real bugs I hit building it, because "here's what broke and how I found it" is a lot more honest than a highlight reel.


What it actually is

CareerMail is an event-driven AI agent, not a chatbot. The distinction matters: there's no UI, no session, no "conversation" in the traditional sense. It's a Gmail inbox with a real pipeline wired behind it:

Email arrives โ†’ Gmail detects it โ†’ Pub/Sub notification โ†’
Cloud Run wakes up โ†’ 7-node LangGraph pipeline โ†’ reply sent โ†’
everything logged to Postgres

The whole thing runs on Cloud Run, Cloud SQL, Pub/Sub, and Secret Manager, with Pinecone + Voyage AI for retrieval and Claude for reasoning and drafting. It costs a few dollars a month to run, and between emails, nothing is running at all โ€” Cloud Run scales to zero.

This is actually my second version of this architecture. The first was a customer-support agent for a fictional client, which taught me the fundamentals โ€” event-driven ingestion, RAG, guardrails, deployment. CareerMail is the second pass, built specifically to demonstrate the pattern on myself, in public, with a genuinely public inbox anyone can test.


The architecture, briefly

Seven nodes, each with a narrow, single job:

  1. load_memory โ€” pulls recent conversation history with this sender from Postgres, so a short follow-up email makes sense in context

  2. question_router โ€” classifies the email, and if it's a genuine inquiry, decomposes it into one or more distinct questions, each routed to the right knowledge topic

  3. retrieve โ€” embeds each question independently and searches only its assigned topic

  4. draft โ€” writes the actual reply

  5. guardrail โ€” an independent safety check before anything gets sent

  6. send โ€” replies via the Gmail API, correctly threaded into the original conversation

  7. log โ€” writes the full outcome to Postgres, on every path, not just successful ones

The full diagram (hand-drawn in Excalidraw, then refined over two passes) is below.

A few of these decisions are worth explaining, because they weren't obvious going in.

Multiple topics, one Pinecone index

I wanted the agent to pull from distinct knowledge areas โ€” career goals, background, technical skills, rates, past work โ€” rather than one undifferentiated blob of "facts about me." The natural approach is separate vector indexes per topic. Pinecone's free tier caps you at five serverless indexes per project, and I'd already used some of that quota on my first project.

The fix ended up being the better design anyway: one index, five namespaces. Namespaces in Pinecone are a hard isolation boundary โ€” a query scoped to rates-availability genuinely never sees about-me content, even if it would score higher by coincidence. Functionally identical to five separate indexes for routing purposes, without the quota cost.

Decomposing compound questions

A real email might ask three unrelated things in one message โ€” "what's your rate, do you know LangGraph, and are you available this month?" Embedding that whole email as one query and searching once produces mediocre retrieval for all three questions, since they get blended into a single, muddy vector.

So question_router doesn't just classify โ€” it splits a compound email into separate questions, each tagged with the single most relevant namespace to search:

class QuestionRoute(TypedDict):
    question: str
    namespace: Literal["career-goals", "about-me",              "technical-skills", "rates-availability", "work-experience"]

Each question gets embedded and searched independently, and the draft step weaves the retrieved facts back into one coherent reply โ€” not a rigid Q1/Q2/Q3 list.

An Isolated guardrail - deliberately bind to the original email

This is the part I think about the most. Before any reply goes out, a separate model call reviews it โ€” but it only ever sees the drafted reply and the distilled questions, never the original email that triggered it.

Why? Because if a prompt injection attempt in the original email successfully influenced the draft, re-exposing that same adversarial text to the "safety check" risks the exact same manipulation succeeding twice. By only judging the output, never re-reading the input, a successful injection has to make the drafted reply itself look suspicious to get caught โ€” a much narrower target than fooling a check that's re-reading the original attack.

Two AI agents drafting together, on purpose

The first version's drafting step was a single model call. The formatting was fine, but not great โ€” a bit list-y, occasionally missed nuance in compound questions. For v2, I used AutoGen to build a small two-agent conversation: a planner that outlines how to cover everything cleanly, and a writer that actually produces the email, with the planner reviewing and either approving or sending it back for another pass.

Both agents see exactly what the guardrail sees โ€” questions and retrieved facts, never the raw email โ€” for the same isolation reasoning. It costs more per email (multiple model calls instead of one), and I think that tradeoff was worth it specifically for output quality; whether it's worth it for every project is a genuine "it depends."


The bugs, because this is the actually useful part

Anyone can show you a working demo. Here's what I actually broke building it.

A guardrail gap, not a guardrail failure. Early on, a drafted reply went out signed "Best, [Your Name]" โ€” a literal, unfilled template placeholder, sent to a real inbox. The guardrail passed it. Looking closer, the guardrail's checks were things like "stays on-topic," "no unsupported claims," "no leaked instructions" โ€” and genuinely none of them covered "contains an obviously unfinished placeholder." The check was doing exactly what it was told; it just wasn't told to check for this. Fixed by telling the writer agent my actual name directly (so it never needed to guess), and adding an explicit placeholder check to the guardrail as real defense-in-depth.

The AI describing itself got flagged as a security risk โ€” correctly. One of my knowledge chunks mentions that this very agent is one of my projects โ€” a nice self-demonstrating detail. The drafted reply phrased it as "this very email agent you're talking to," and the guardrail blocked it: "This blurs the line between a human professional reply and disclosure of system internals." That's genuinely sound reasoning โ€” a reply describing its own architecture in first person is a real red flag in most contexts. It just needed an explicit, narrow exception for this one legitimate case, plus a content rewrite so it read as "I built this," not "I am this."

A real inquiry got silently escalated. Someone emailed asking about my tech stack โ€” genuinely answerable โ€” and also asked me to scope a hypothetical inventory-management AI system โ€” not something I can answer from stored facts. The system correctly recognized it had no basis to answer the second part and escalated, which is the right call architecturally... except escalation currently means total silence to the sender. From their side, that looks identical to the system being broken. That's now a concrete, planned v2 feature: partial answers for what's genuinely answerable, plus a proposed-meeting flow (tentative calendar hold, human-confirmed, not autonomous booking) for anything that needs a real conversation. Not something that I can assure will 100% be there, but its just a thought.

Logs that silently weren't there. After deploying, requests were returning success, but nothing was showing up in Cloud Run's logs โ€” not even delayed, just gone. Turned out to be Python's stdout buffering: inside a container, with output piped rather than attached to a real terminal, print() statements can sit in an internal buffer indefinitely instead of actually reaching the log stream. One line, ENV PYTHONUNBUFFERED=1, fixed it permanently โ€” a good reminder that "my code has no bugs" and "my code's output is invisible" can look identical from the outside.

Building on Apple Silicon for a cloud that isn't. Docker on an M-series Mac builds images for arm64 by default. Cloud Run runs amd64. The fix is a single flag (--platform linux/amd64), but the failure mode โ€” a deploy that silently accepts an incompatible image and then fails at runtime โ€” is exactly the kind of thing worth knowing about before it costs you an afternoon.


What's deliberately not built yet

Scope discipline matters as much as what you ship. v1 is intentionally text-only: no PDF or image attachments read from incoming mail, no resume sent as an attachment, no calendar integration. Each of those is a real, considered v2 item, not a forgotten gap โ€” writing them down explicitly, rather than bolting them on halfway through, is part of what kept this buildable in the time I had.


Try it

If any of this was interesting, the best way to actually evaluate it isn't this post โ€” it's emailing askrajat.ai@gmail.com directly. Ask it something real. Ask it something adversarial, if you're curious how the guardrail holds up. It's a genuinely live, working pipeline, not a static demo, and I'd rather you form your own opinion of it than take mine.

My Projects

Part 1 of 3

In this series I will be showcasing some of the projects that I have been working on. In this series I will be sharing some of the architectural designs that I have designed. I am open to suggestions but do keep in mind I am about to become a fresh graduate

Up next

Building an Agentic Job Searcher โ€” The Low Level Design

This is the second post in my series about building an agentic AI system that searches for jobs, tailors resumes, and applies automatically based on my resume and preferences. If you haven't read the