The problem
A founder can buy a list of companies for fifty dollars a month. What they can't buy is the ten minutes of research between that list and the first email.
That research is the job: the one specific thing worth opening with, and proof that it's true. A founder's real fear isn't a slow pipeline. It's sending something wrong under their own name.
Two limits came with it:
- Nothing gets sent
- No personal contact data
What I asked first
The brief left these open. I answered each one before writing any code.
- How agentic, exactly? One long run wasn't needed, and the hosting's time limits rule it out anyway. So it's four short stages. I set the stages. The agent picks its tools inside them.
- What if the evidence doesn't support ten? Fewer strong leads beat a longer weak list. So it widens the search and never lowers the bar. Still short, it stops and asks the founder.
- Ask the founder, or refine silently? Both. It turns the goal into criteria, marks each one hard or soft, and shows them before a single paid search.
- What does approval mean when nothing is sent? A gate before sending would be theatre. So the gate is on spending: the founder confirms the criteria before any money moves.
- What is the budget for? The search account is shared, so overspending takes from other people. Every run has hard caps on companies, searches, page reads and spend, with an estimate up front.
How it runs
- Person
- AI
- Tool
- Check
- Result
Brief
01, Human:
Founder sets the goal
Founder
The kind of company worth talking to.
02, AI:
Criteria written
Agent
Each one marked hard or soft.
03, Human:
Founder confirms
Founder
Nothing paid has run yet.
Discover
04, Step:
Find companies
Apify
Hard caps on searches and spend.
Qualify
05, Step:
Read each site
Firecrawl
Contact details removed first.
06, AI:
Qualify with evidence
Sonnet 5
Every verdict cites what it read.
07, Logic:
Enough leads?
Check
Still short after the refills, it asks the founder.
Then: step 4 (short? widen the search); step 8 (yes).
Draft
08, AI:
Draft outreach
Sonnet 5
3 emails and a LinkedIn note per lead.
09, Output:
Saved, not sent
In the app
There's no send tool.
Enforced, not requested
- No tool can break scope There's no email finder and no send tool.
- Contact data never lands Contact details are removed before anything is stored.
- Every claim has a source A claim that points to something it never read is rejected.
- The bar never drops The target is reached by widening the search.
- Every stage has a limit A cap on how many steps it can take and how much it can spend.
- No borrowed connectors It can't pick up the apps connected on my own machine.
- Everything is logged Every action, including the ones that were blocked.
The thing that nearly got past me
IncidentLogged 24 September 2026
An agent that must never send had an email connector attached
Until then, every stage ran with my own machine's setup: my settings, my notes and my connected apps, Gmail included. It ran on my personal account and paid to read all of that at every stage. The whole promise of this build is that nothing gets sent, and an email connector was attached to every stage.
The fixThe agent now runs in its own folder with its own settings: only its own tools, a fixed set per stage, no access to files, and its own account. The apps connected on my machine can't get in.
The run, in numbers
12
qualified leads, against a target of 10
48
outreach drafts, saved for a person. None sent.
$0.137
spent on finding companies
72
actions the agent took, every one logged
0
actions blocked for going over a limit
2
errors it recovered from
$1.996
estimated AI cost, kept apart from real spend
All from one full run of the agent.
What it refuses to do
- Send anything Drafts are saved for a person to use.
- Collect contact details They're removed before anything is stored.
- Lower the bar It widens the search to hit the target instead.
- Make a claim it can't back up Every claim points to something it read.
What I took from it
A rule the agent can choose to ignore isn't a rule. The ones that matter here are built in, so they hold even when it ignores them.
The limit I'd fix first
if I did it again
- A lead's "why now" is optional, and nothing checks it. Only 3 of the 12 qualified leads had a real timing signal.
- The step that writes the criteria sets the rules for everything after it, and nothing checks its work. It came back with no deal-breakers at all, and all seven later rejections belonged on that list.
