Skip to content

About

Develop an AI agent that generates a daily list of suitable clients for Open Numerics and Reaction Studio

Resources

Stars

2 stars

Watchers

0 watching

Forks

Repository files navigation

Prospectus-Agent

A little agent that does the part of running a business I'm worst at: finding people to email.

I run Open Numerics, a small scientific-computing consultancy. We're good at the work, not at prospecting — every week I'd lose a morning opening fifteen tabs, fire off three awkward cold emails, and never follow up. So I built this: each morning one command finds companies that fit, writes a tailored email for each, and — once I've looked at the dry run — sends them from my own Gmail and follows up when a thread goes quiet. The worst hour of my week became a few minutes over coffee.

Nothing in it is specific to me — the business-specific bits live in one config file, so you can point it at your own. (It ships configured for Open Numerics; swap in yours.)

What it does

  • Finds fitting companies — searches the web each run, scores for fit, and skips anyone you've already seen. No bloated lead list to pay for.
  • Reads before writing — the draft talks about their actual work, in a plain human voice (no mail-merge filler, no AI tells).
  • Finds the right person's real email, not a generic info@ — checks the contact page and footer, infers the company's address format from a real address and applies it, and optionally verifies deliverability. One best address per senior person.
  • Sends them from your Gmail, or hands you copy-paste-ready drafts — your choice. Sending is a dry run by default and only goes live with an explicit --live.
  • Follows up when a thread goes quiet (that's where most replies come from), threaded under the original email so it lands in the same conversation. Two nudges, then it stops.
  • Runs on a cheap model for pennies a day, not a monthly sales-SaaS seat.

Make it yours

You don't edit code — you write one file, profile.yaml: what you sell, who's a good fit, who to exclude (competitors), which industries to focus on, which giants to avoid, and how your emails should sound (voice_notes, example_openers, credibility, opening_style). See profile.example.yaml for a worked example.

python3 -m venv .venv
.venv/bin/pip install -e ".[dev]"       # package + the `prospectus-agent` command
cp .env.example .env                    # add your ANTHROPIC_API_KEY
cp profile.example.yaml profile.yaml    # describe your business + ideal customer

That's the whole setup (if profile.yaml is missing, the example is used). The install puts prospectus-agent on your PATH and resolves its files relative to the project folder, so you can run it from any directory — point PROSPECTUS_AGENT_HOME elsewhere to relocate them. Prompt templates live in src/prospectus_agent/prompts/; edit only to change wording/tone.

Two extras are opt-in: pip install -e ".[gmail]" for live sending (dry runs need nothing), and pip install openai if you point a role at OpenAI.

More than one business: a profile named acme lives in profile.acme.yaml with its own acme.db, outbox/acme/, and brief cache — fully isolated. Run prospectus-agent --profile acme, or set DEFAULT_PROFILE=acme in .env. Each profile carries its own voice (capability_areas, voice_notes, credibility, opener examples, recent_innovations), so the engine stays business-agnostic.

How it works

prospectus-agent runs the pipeline:

  1. Refresh profile — fetches your website (cached weekly via PROFILE_REFRESH_DAYS) so drafts reflect what you do now.
  2. Discover — up to MAX_DISCOVERY_CALLS web-search rounds for companies scoring ≥ FIT_SCORE_THRESHOLD/10, rotating the industry angle and skipping anything already seen. Every qualified, reachable company found is drafted — there's no daily target and no backlog; a run yields however many were discovered, and --deliver sends them all. Everything seen is recorded either way (not_a_fit for competitors/oversized/excluded sectors, unreachable for dead domains), so nothing resurfaces.
  3. Research + draft — two calls per company: the searcher reads the site (contact page + footer, where addresses live) and leadership and returns grounded facts; the writer turns those into the email. Contacts are one best address per senior person: a published one if found, else the domain's inferred format applied, else a first.last@ guess. Generic info@ is a fallback only; dead domains (no MX) are skipped; opt in per profile to verify addresses via Verifalia (catch-all aware, fails open). The email is drafted in a human voice — no em dashes or AI clichés, subject led by your company name, no signature (Gmail's own signature is appended on send).
  4. Follow-ups — flags sent companies with no reply after FOLLOWUP_DAYS (default 5) calendar days; two nudges max (a fuller first, weaving in a recent_innovations win, then a short final touch-base), then the company goes terminal (no_reply).
  5. Writes outbox/<date>/ — new_prospects.{md,html} + followups.{md,html}, each with its contact list and a copyable comma-separated To: line. The HTML links your company name so pasting into Gmail keeps the hyperlink; re-running the same day appends rather than overwrites. This is your review surface whether or not you let the agent send.
  6. Prints a digest, a token-usage line, and an estimated dollar cost per model.

prospectus-agent --refine re-drafts today's existing drafts with the current prompt (no re-discovery or research — cheap and fast), then regenerates the outbox; curated contacts are left untouched.

Social posts (--socialmedia)

A separate, opt-in scope: three drafted posts a week, generated together so you can schedule them in advance. Monday teaches one technical thing, Wednesday walks through a real project (often as an instalment of One Computational Bottleneck — problem, why it mattered, how it was solved, the lesson), and Friday is a founder opinion. Each comes with a LinkedIn version and an X version: a numbered thread for Mon/Wed, a single post for Friday.

Raw material comes from your own site (company.website_path, including every case study under solutions/) plus the profile's recent_innovations. Anything the post claims about work you've done has to trace back to that material — the writer is told to make a general technical argument rather than attribute an invented detail to a project.

Two things it does that a plain "write me a post" wouldn't. Friday may return a skip recommendation instead of a confident post: a bland opinion costs more credibility than silence, so the draft is still written but flagged. And every X part is length-checked in code, because a model cannot reliably count characters — over-length parts are named in the digest and in the outbox rather than silently shipped.

Enable per profile with settings.social_media: true; it's off by default, since a voice built for one business doesn't transfer to another. Drafts land in outbox/<date>/social.{md,html} and in the posts table. Nothing is ever published — --socialmedia writes, you edit and post. --socialmedia --refine re-drafts that same week under the current prompt, keeping each post's topic so you can judge a prompt change rather than a new subject.

Sending (--deliver)

--deliver sends the queue — drafted initials plus any due follow-ups — through the Gmail API, so copies land in your Sent folder. It uses no LLM: it reads drafts and contacts, sends, records the send, and starts/advances the follow-up clock.

  • Dry run unless you pass --live. A dry run prints, per company, exactly who it would write to (each address tagged verified/public/inferred/guessed), the subject, and whether a follow-up would thread. Nothing is touched.
  • Recipients: personal published/verified/inferred addresses go in To, along with a public generic inbox; guessed personal addresses go in Bcc, and only when no confirmed personal address exists. If To would contain only a generic inbox, the best personal address is promoted into it — the goal is always to reach a human. Capped at AUTOSEND_MAX_RECIPIENTS (lowest confidence dropped first, never the last person).
  • Threading without inbox access: each send gets a self-owned Message-ID on your domain, stored with Gmail's message/thread ids. A follow-up reuses them (Re: …, In-Reply-To, References, same threadId), so it lands in the original conversation — no inbox read scope needed. With no stored thread id it goes out as a fresh email instead.
  • Two safety gates: the profile must be listed in AUTOSEND_PROFILES, and all three GMAIL_* credentials must be set. Fail either and --live downgrades to a dry run with a notice rather than doing something surprising. Companies with no deliverable address are skipped and listed by name.
  • Sends are paced AUTOSEND_PACING_SECONDS apart, and the DB is snapshotted to <db>.bak before any mutating run.

One-time Gmail setup: create an OAuth Desktop client for an Internal app (Google Cloud Console → APIs & Services → Credentials), then

.venv/bin/pip install -e ".[gmail]"
.venv/bin/python -m prospectus_agent.gmail_auth --client-id XXX --client-secret YYY

It opens a browser for consent and prints the three GMAIL_* values to paste into .env. Scopes are gmail.send and gmail.settings.basic (to read your signature) — no read access to your mail.

Two roles, your choice of model — and vendor. A cheap searcher (discovery, research, profile refresh, with web_search; default Anthropic claude-haiku-4-5) does the high-volume work; a stronger writer (every email; default claude-sonnet-4-6) writes the short drafts. Mix Anthropic and OpenAI freely in .env — both go through one vendor-neutral llm.py seam (Messages / Responses API, structured output via strict tools).

Config dials (.env)

profile.yaml is your business; .env is the machinery (gitignored). Anything here can also be overridden per business via a settings: block in that profile's YAML (profile settings win over .env).

Variable Purpose
ANTHROPIC_API_KEY / OPENAI_API_KEY Key(s) for the vendor(s) you use — only those named in SEARCH_VENDOR / WRITER_VENDOR.
DEFAULT_PROFILE Business to run with no --profile (loads profile.<name>.yaml + <name>.db + outbox/<name>/).
SEARCH_VENDOR / SEARCH_MODEL Vendor (anthropic|openai) + model for the searcher. Default anthropic / claude-haiku-4-5.
WRITER_VENDOR / WRITER_MODEL Vendor + model for the writer (drafting). Default anthropic / claude-sonnet-4-6.
DISCOVERY_MODEL Override the searcher model for the mechanical steps only; defaults to SEARCH_MODEL.
FIT_SCORE_THRESHOLD, MAX_DISCOVERY_CALLS, FOLLOWUP_DAYS Pipeline tunables (FOLLOWUP_DAYS = calendar days before a follow-up is due, default 5). Every qualified, reachable company discovered is drafted — no daily target.
TARGET_REGION Fallback geographic focus; prefer targeting.region per profile.
AVOID_SECTORS Comma-separated sector keys to exclude entirely (out of scope → not_a_fit). Valid keys in sectors.py.
MAX_COMPANY_SIZE Largest size to target: startup|small|mid|large|enterprise (default mid).
MAX_PUBLIC_EMAILS / MAX_PEOPLE / GUESSES_PER_PERSON Contact-list size (≤3 people, one address each; inbox as fallback). Keep GUESSES_PER_PERSON=1.
AUTOSEND_FROM The From address for --deliver (must be the Gmail account you authorized).
AUTOSEND_PROFILES Comma-separated profiles allowed to send; everything else is dry-run only.
AUTOSEND_MAX_RECIPIENTS / AUTOSEND_PACING_SECONDS To+Bcc cap per email (default 5) and the gap between live sends (default 2s).
GMAIL_CLIENT_ID / GMAIL_CLIENT_SECRET / GMAIL_REFRESH_TOKEN Gmail OAuth creds from python -m prospectus_agent.gmail_auth. All three required for --live.
VERIFALIA_USERNAME / VERIFALIA_PASSWORD Optional Verifalia HTTP-Basic creds for mailbox verification. Blank = off. Enable per business with settings.verify_emails: true in its profile.<name>.yaml.
VERIFY_MAX_CANDIDATES Max address formats verified per person (default 2 — keeps a 25/day free tier in budget).
DENY_LIST_LIMIT, PROFILE_REFRESH_DAYS Size of the "don't repeat" hint sent to the model (the DB still dedups fully), and how often your own profile is re-fetched.
DISCOVERY_EFFORT / DRAFTING_EFFORT / SEARCH_CONTEXT_SIZE OpenAI backend only; ignored on Anthropic.
DISCOVERY_MAX_TOKENS / DRAFT_MAX_TOKENS / PROFILE_MAX_TOKENS Per-step output-token caps.

Daily use

The CLI is scope × action. Scope: which business (--profile, or --runall for all of them) and which drafts — new prospects by default, follow-ups with --followup, social posts with --socialmedia. Action: --refine (re-draft), --deliver (send), or --sent (record that you sent them by hand); no action = discover + draft. The three actions are mutually exclusive, and the CLI tells you when a combination isn't meaningful.

prospectus-agent                    # NEW PROSPECTS: discover + draft (+ sweep follow-ups)
prospectus-agent --refine           # re-draft today's with the latest prompt

prospectus-agent --followup           # FOLLOW-UPS only: draft one for anyone past the threshold
prospectus-agent --followup --refine  # re-draft them with the latest voice

prospectus-agent --deliver            # SEND: dry run over every profile — see what would go out
prospectus-agent --deliver --live     # actually send initials + due follow-ups, all profiles
prospectus-agent --profile acme --deliver --live   # just that business
prospectus-agent --followup --deliver --live       # due follow-ups only

prospectus-agent --socialmedia           # SOCIAL: draft next week's Mon/Wed/Fri posts
prospectus-agent --socialmedia --refine  # re-draft that week with the latest prompt

prospectus-agent --profile acme       # any of the above, for a different business
prospectus-agent --runall             # daily pipeline for EVERY profile (forwards action flags)

prospectus-status drafts            # list drafts ready to review
prospectus-status show DOMAIN       # full draft + contacts
prospectus-status mark DOMAIN sent  # (or: replied / not_interested) — drives the follow-up clock

--deliver without --profile fans out across every profile, so one command sends the whole day's queue. Sending by hand instead? Copy from outbox/<date>/*.html and then run prospectus-agent --sent (or --followup --sent) to start the follow-up clock — the manual path still works exactly as before.

(Activate the venv, or prefix with .venv/bin/.) Status values: new, drafted, sent, followed_up, no_reply (both follow-ups sent, done), replied, not_interested, not_a_fit, unreachable. Replies are tracked manually — the agent has send-only Gmail access and never reads your inbox, so prospectus-status mark DOMAIN replied is what stops the follow-ups.

A few honest notes

  • Dry-run first, always. --deliver without --live is free and tells you exactly what would go out. Read the drafts in outbox/ before you hand it the keys.
  • Guessed emails are guesses. Unverified ones may bounce; verify (or turn on Verifalia) for anyone important. The dry run tags every address with how it was derived.
  • Garbage in, garbage out. Lead quality tracks how well you describe your ideal customer in profile.yaml. Spend ten minutes on it.
  • Watch the cost line. The defaults split the work (cheap searcher, stronger writer); the token + estimated-dollar summary each run is there so you notice early.
  • Cold email is your reputation. Volume is bounded only by MAX_DISCOVERY_CALLS, so ramp up slowly and keep the writing worth reading.
  • Single-user and local (a SQLite file per business, your own Gmail). A hosted multi-tenant version is a someday-maybe.

The code

Small and deliberately boring. Engine code lives in src/prospectus_agent/ and knows nothing about any business — everything company-specific is in profile.yaml and prompts/.

Area Files
CLI / entrypoints cli.py (flag validation + dispatch), daily_run.py, refine.py, followup_run.py, deliver_run.py, mark_sent.py, social_run.py (--socialmedia), mailbox_run.py (--bounces / --replies), run_all.py (--runall fan-out), status.py
Config / setup paths.py (home + .env), config.py (.env + profile settings: + per-profile paths + client factory), runner.py (session prologue + DB backup), agent_profile.py (loads profile.*.yaml)
Pipeline discovery.py (+ sectors.py sector classifier for AVOID_SECTORS), research.py (research + initial draft), followups.py (follow-up sweep + generator), redraft.py (--refine), on_profile.py (website brief), site.py (local website reader), social.py (weekly posts)
Contacts / email contacts.py (address pattern inference + guessing), verify.py (MX + Verifalia), outbox.py (md/html digests), send.py (recipient policy, MIME + threading, Gmail seam), gmail_auth.py (one-time OAuth), gmail_read.py (paced batch reads), bounces.py (NDR parsing), replies.py (reply detection)
Data / model db.py (SQLite), llm.py (vendor-neutral seam + usage/cost), prompts/, schemas.py

The delivery policy is written up in docs/auto-send-design.md.

Tests

.venv/bin/pip install -e ".[dev]"
.venv/bin/python -m pytest

212 tests, all offline — both vendor clients are faked and the Gmail service is stubbed, so no key, no network, no spend. Covers the database, address inference/verification, follow-up timing and the follow-up state machine, the discovery loop (competitor + over-size + unreachable filtering), research/drafting, recipient selection and message threading, dry-run vs live delivery, the LLM helpers, and the strict tool schemas.

Roadmap / known gaps

  • Follow-up timing is plain calendar days (FOLLOWUP_DAYS); no working-day/holiday calendar.
  • Replies are marked by hand — the agent is send-only by design and can't detect a reply itself.
  • Address verification is opt-in; unverified guesses can still bounce.
  • Single-tenant and local; a configurable hosted version may come later.

Contribute

Contributions very welcome — this is meant to be shared. Lots of businesses face the same prospecting struggle; let's help each other.

About

Develop an AI agent that generates a daily list of suitable clients for Open Numerics and Reaction Studio

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages