Project Info
Inspiration
We didn't start with a piece of technology we wanted to use. We started with a moment we kept hearing about, over and over, when we talked to parents of autistic kids. The moment right after the diagnosis appointment, standing in a parking lot, holding a folder full of paper, and realizing nobody had told them what to actually do next. Not what's wrong. They already knew that. What to do. Which specialist. Which form. Which benefit they probably qualify for but will never hear about unless they happen to ask the right person the right question. Who to call at nine at night when something's hard and they just need another human being to talk it through with, and the only humans available are exhausted too. That parking lot moment is what Compass is for. Not the diagnosis. The everything after.
What it does
Compass is a navigator, not a database. A parent uploads an IEP, an evaluation, a therapy note. And Claude reads it the way an exhausted parent at midnight can't, and turns it into a timeline and a short, clear list of what to do next. Need a speech therapist who actually takes your insurance? A specialist search agent goes and looks, live, right now, instead of handing back a directory that's been stale since 2022. Wondering if your state has a waiver program you've never heard of? Same thing, a real search, not a guess. And then there's the part of the app we care about most, because it's the part that shows up at the worst moments: a single mic button. No form to fill out while you're upset. Press it, and an AI comfort companion is immediately there, talking with you, present with you, while you figure out what's going on with your kid. It is never pretending to be a doctor. It never promises something it can't know. It is just there, which is sometimes the entire thing a person needs at 9pm. From there, Compass also opens a request for a real human to join, right now, today that means a member of the general public through Terac's annotator pool, and we say that plainly rather than dress it up as something it isn't. The path forward, and the thing we actually want to build next, is a version of that same human layer screened and verified for medical professionals, so the person who picks up is exactly who a parent needs in that moment. What makes the AI companion trustworthy isn't just that we wrote it a calm system prompt. It's that real people, through Terac, listen to what it produced and tell us, honestly, whether it was actually good. And we used that feedback to make it better, and then proved it got better, blind, with a second round of real people judging the before and after without knowing which was which.
How we built it
We built Compass as a Next.js app on Supabase, with Claude doing the reading, summarizing, and reasoning throughout. Extracting structure out of messy uploaded documents, generating next steps, moderating a community feed, writing plain English explanations of dense IEP language. For the parts of the app that needed to know things about the world, not just reason over text, we used Browserbase and Stagehand to run real browser agents against Psychology Today and findhelp.org, so the specialist and benefits search results are genuinely current, not something we faked to look alive. The live connect feature, the mic button, is where the architecture got genuinely complicated, in a good way. A press of the mic spins up a private Daily.co video room and starts a Vapi voice agent talking to the parent, live, in the browser. When the conversation ends, Vapi hands us a summary, and that's when we reach for Terac. We create and launch a real opportunity asking a general population annotator to read that summary and tell us, honestly, how clear and helpful it actually was, and to rewrite it if it wasn't good enough. That correction doesn't just sit in a database. We built a four script pipeline that turns real corrections into an improved version of the summarizing prompt, runs both versions against held out test scenarios, and then asks a second, separate group of Terac annotators to blindly judge which one is better, with no idea which prompt produced which answer. That blind judgment is what let us say, honestly, that human feedback measurably made the system better, not because we believed it would, but because we watched real people prefer it, 8 times out of 8. A teammate built the Poke integration in parallel on its own branch, a parent can text Compass related questions over iMessage without ever opening the app, and we merged our work together once both pieces were solid.
Challenges we ran into
Almost nothing about this went in a straight line, and we think that's worth saying honestly instead of pretending it did. We started building toward connecting parents with real, credentialed specialists through Terac, before realizing, partway through a working integration, that the hackathon's actual rules require general population annotators only, no clinical panels. That meant throwing out targeting logic we'd already built and rethinking what the feature was actually for. It also meant being honest in our own product about who a parent is actually being connected to right now, a general member of the public, not a verified professional, and treating that as a clearly labeled current limitation rather than something to gloss over. We hit a real, documented outage on Terac's launch endpoint mid build. We proved it wasn't our bug by replaying a byte for byte successful payload and watching it fail identically, and had to keep building around a dependency we didn't control. We accidentally left two test opportunities running and burned real credits before we caught it, which taught us fast to put a hard cost cap in code, not just in our heads. We learned the hard way that a JWT with dots in it will break a URL templating system on the other end, and that the fix that came out of that bug, never handing a sensitive token to a third party at all, routing through our own redirect instead, ended up being more secure than what we'd originally designed.
Accomplishments we're proud of
We're proud that the human in the loop story isn't a slide, it's real data. Real annotators, real corrections, a real blind comparison, an actual 8 out of 8 preference result we can show, not just claim. We're proud of the mic button. It sounds small. It isn't. Replacing a form with a single press, and trusting an AI conversation to gather what a form would have demanded up front, is a genuine bet on meeting someone where they are instead of making them perform organization while they're scared. We're also proud that we held the line on honesty about who a parent is actually talking to right now. It would have been easy to quietly imply the person on the other end of that call was a credentialed professional. We didn't, and we built the screened, verified version of that into our roadmap instead of our marketing. And we're proud that when we hit real, unglamorous problems, an outage, a billing scare, a confusing third party bug, we slowed down, diagnosed them properly instead of guessing, and fixed the actual cause instead of papering over the symptom.
What we learned
We learned that "use the human in the loop" sounds like a feature checkbox until you actually build it, and then it becomes a real design question. Who are these humans, what are you asking them to judge, and how do you know their feedback actually made anything better rather than just feeling like it did. The blind comparison wasn't a nice to have. It was the only honest way to answer that question. We learned that reading a sponsor's actual rules document, closely, more than once, costs you an hour and saves you a weekend. We learned that the gap between "general population reviewer" and "verified medical professional" isn't a wording choice, it's a real safety boundary, and that the right move when you haven't built the second one yet is to say so clearly, not to describe the first one as if it were the second. And we learned, again, that the parents we built this for don't need another app that's impressive. They need one thing to actually work when they're tired and scared and it's late. That's a much higher bar than it sounds like, and we tried to hold ourselves to it.
What's next
The most important next step is the one we already named honestly in this README. Taking the general population annotator pool we used for this hackathon and building the real version on top of it, a screened, verified pathway where the person who joins a parent's call is confirmed to be a medical professional, with real credential checks, not just an open task on a marketplace. That's not a someday idea, it's the direct next milestone for live connect. We'd also want to run the before and after evaluation at real scale, not 8 scenarios, enough comparisons to trust the number, not just be encouraged by it. And we'd want to close the loop further, letting the live connect companion get better continuously, from every real call, not just the ones we happened to run during a weekend. Mostly, though, what's next is the same thing that was first. Finding the next parent standing in a parking lot with a folder full of paper, and making sure Compass is the thing that tells them what to do next, and eventually, who to safely talk to.
Compass
Compass is a navigator app for parents of autistic and special-needs children. Getting a diagnosis is the easy part — the hard part is everything after: figuring out what an IEP actually means, finding a therapist who takes your insurance, learning what your state will actually pay for, and not feeling alone while you do it. Compass puts all of that in one place and uses AI to do the research a parent would otherwise spend hours doing themselves.
This README covers what's actually built and working today, including what was built specifically for the Terac challenge.
Hackathon Tracks
Compass was built for, and submits to, three tracks:
- Terac — the live-connect annotation + blind A/B evaluation flow described in detail below (real human annotators rating and correcting AI summaries, then a second blind comparison measuring whether that feedback actually improved the model).
- Browserbase — the Directory (specialist search) and Benefit Finder features both run live Browserbase/Stagehand agents against real sites (Psychology Today, findhelp.org) rather than static or mocked data.
- Poke — the onboarding/Settings phone-number handoff that connects a parent to a pre-built Poke recipe for texting Compass-related questions.
Key Features
Roadmap
Upload an IEP, evaluation, or therapy note (PDF or image) and Claude extracts diagnoses, current services, goals, recommendations, and important dates directly from the document. Those extracted items build a timeline, and a separate Claude call generates 3–6 concrete "next steps" (with urgency: now / soon / upcoming) based on the child's profile and extraction history. This is the entry point for a parent who just got a new document and doesn't know what to do with it.
Compass Coach
Upload a document and get a plain-English, section-by-section breakdown: what each part means, which parts are worth raising concerns about, and a list of questions to bring to the next IEP meeting/doctor appointment. Stateless — nothing is persisted, it's a one-time analysis tool.
Directory (Specialist Search)
Search for ABA, speech, OT, PT, psychology, or developmental-pediatrics providers by ZIP code. Live results come from a Browserbase/Stagehand agent that searches Psychology Today and extracts real listings (name, phone, address, description, profile link), cached for 7 days per ZIP+specialty. Results can be saved for later, and each one can get an on-demand Claude-generated summary.
Benefit Finder
Same Browserbase/Stagehand approach, pointed at findhelp.org, searching for Medicaid waivers, SSI, Regional Center / state DD-agency programs, and other disability-specific assistance — targeted using the child's diagnosis, age bracket, and current services so results lean toward specific, relevant programs rather than generic disability listings. Results are cached for 7 days and can be saved.
Village
A public community feed — parents post under topics (newly diagnosed, IEP help, school, behavior, therapies, general) and reply in threaded comments. Every post and comment passes through a Claude moderation check before publishing. Deliberately feed-only: no private messaging between parents.
Poke (text assistant)
During onboarding (or later in Settings), a parent can save their phone number and connect to a pre-built Poke recipe, letting them text Compass-related questions via iMessage/SMS without opening the app. The phone number is the only thing Compass itself stores and manages — the text-handling logic lives in Poke's hosted recipe.
Live-Connect (mic-first AI intake)
This is the feature built specifically around the Terac challenge — see the next section for the full breakdown of how it works and what it measures.
The Terac Challenge: How Compass Uses It
The core idea: instead of routing a parent to a real human specialist (a licensed clinician) for an emotionally difficult moment, Compass routes them to an AI comfort companion first, and uses Terac's general-population annotator marketplace to evaluate and improve the AI's output with real human judgment — without needing licensed clinicians in the loop at all.
The flow
- A parent presses a single mic button on
/connect. No form, no waiting room. - A private Daily.co video room is created immediately for that request (in case anyone wants to join later).
- A Vapi voice agent talks to the parent in-browser — a "comfort companion," not a clinician — and listens to whatever they want to talk through.
- When the call ends, Vapi generates a summary and posts it to Compass's webhook.
- Compass creates and launches a Terac opportunity: a task asking a general-population annotator to read the AI-generated summary and rate it.
- Compass polls Terac for a claimed submission. Once someone claims the
task, it's auto-approved and the request status flips to
scheduled. The parent gets a link to join the Daily call room if they want to talk live to whoever picked up the task; the annotator gets the same link from their task page.
What the annotator actually does
The Terac participant lands on a public Compass page
(/annotate/{requestId} — no login required, since they have no Compass
account) and:
- Reads the AI-generated summary of the parent's call.
- Rates its clarity on a 1–5 scale.
- Optionally writes a corrected/improved version of the summary.
- Optionally leaves free-text notes.
That submission is stored in an annotations table, tied back to the
Terac submission ID. This is the human-data-collection half of the
challenge: real people, not the model itself, judging and correcting AI
output.
The before/after evaluation methodology
The annotation data feeds a second pipeline designed to measure whether human feedback actually improves the AI's summaries — via a blind A/B comparison, also run through Terac, so the evaluation itself is human judgment, not a self-graded metric.
scripts/build-improved-prompt.mjs— pulls real corrected summaries out of theannotationstable and uses them as few-shot examples to build av2_improvedprompt (falls back to hand-written guidelines from common failure patterns if no corrections exist yet).scripts/run-eval.mjs— runs 8 fixed, held-out test scenarios (realistic intake-call transcripts) through bothv1_baselineandv2_improvedprompts via Claude, producing paired summaries.scripts/collect-blind-eval.mjs— randomly assigns each pair's two summaries to slots A/B (the v1/v2 label is hidden), writes them to acomparison_pairstable, and launches a second, separate Terac opportunity — one task per pair, at/compare/{pairId}— asking general-population annotators which summary is clearer, with no idea which one came from which prompt version.scripts/report-results.mjs— talliescomparison_resultsagainst the hidden labels and prints how oftenv2_improvedwas preferred overv1_baseline.
Current status: the annotation pipeline (steps 1–6 of the live-connect
flow above) is live and wired end-to-end through real calls, and the
four-script before/after pipeline has been run for real. v2_improved was
built from 3 real Terac-collected annotation corrections (not synthetic
guidelines), and an 8-scenario blind A/B comparison — also judged by real
Terac annotators, with the v1/v2 label hidden from them — came back
8/8 (100%) in favor of v2_improved. That improved prompt has since
been deployed to the live Vapi assistant (scripts/update-vapi-assistant.mjs
refuses to deploy anything not built from source: "annotations"), so
production calls now use the human-feedback-improved summarizer, not just
the eval script.
Why general-population annotators, not clinicians
Per the Terac challenge's design, annotators are general members of the public, not screened for any clinical credential. Compass's screening question is deliberately minimal (an optional free-text "anything you'd like us to know before starting" field) — there's no profession-based filter. This is a hackathon-rules constraint, not a production design choice: a real deployment of this idea would need either licensed reviewers or a much more careful consent/safety framework before any annotator is shown details from an emotionally vulnerable parent's call.
Tech Stack
Framework
- Next.js 14 (App Router), React 18, TypeScript
- Tailwind CSS
Database & Auth
- Supabase (Postgres, Auth, Row-Level Security, Storage)
AI / LLM
- Anthropic Claude (
claude-sonnet-4-6) — document extraction, IEP analysis, next-steps generation, specialist/benefit summaries, community moderation, and the live-connect eval scripts
Browser automation
- Browserbase + Stagehand — live web search/extraction against Psychology Today (specialists) and findhelp.org (benefits)
Live-connect stack
- Vapi — in-browser voice AI agent (the "comfort companion")
- Daily.co — video call rooms, created per request
- Terac REST API (
https://terac.com/api/external/v2) — opportunity creation, launch, submission polling/approval. Note: a.mcp.jsonpointing at Terac's MCP server exists in the repo as an artifact of early exploration, but the final, working integration is a plain REST client (src/lib/terac/client.ts) — MCP was not used in the shipped architecture.
Other integrations
- Mailgun (HTTP API) — weekly digest emails
- Poke — text-assistant recipe (phone-number handoff only; the conversational logic lives outside this repo)
Architecture
Parent presses mic (/connect)
│
▼
POST /api/connect/request ──► creates expert_call_requests row (pending)
│ creates Daily room + token (src/lib/daily)
▼
Vapi voice agent runs in-browser (ComfortAgentWidget.tsx)
│ parent talks, call ends
▼
POST /api/vapi/webhook
│ stores call_notes.ai_generated_summary
│ creates + launches a Terac opportunity (src/lib/terac/client.ts)
│ task_url → /annotate/{requestId}
│ status: pending → launched
▼
Browser polls GET /api/connect/request/[id]/status every 5s
│ checks Terac submissions; once claimed, auto-approves
│ status: launched → scheduled (or → timed_out after 10 min)
▼
Annotator opens /annotate/{requestId} (public, no auth)
│ rates clarity (1–5), optionally corrects the summary, adds notes
▼
POST /api/annotate/{callNotesId} ──► stores annotations row
│
▼
Parent + annotator can both join the same Daily room via room_url
Eval pipeline (separate, offline, run via scripts/)
│
build-improved-prompt.mjs ──► v2_improved prompt, from real annotations
▼
run-eval.mjs ──► runs 8 held-out scenarios through v1 + v2 via Claude
▼
collect-blind-eval.mjs ──► publishes blind A/B pairs as a second Terac
│ opportunity at /compare/{pairId}
▼
report-results.mjs ──► tallies which version annotators preferred
Every other feature (Roadmap, Directory, Benefits, IEP Coach, Village) follows the same simple pattern: a Next.js page/component calls an internal API route, which calls either Claude directly or a Browserbase/Stagehand agent, and reads/writes Supabase.
Setup Instructions
1. Prerequisites
- Node.js 18+
- A Supabase project
- API keys for: Anthropic, Browserbase, Daily.co, Vapi, Terac, Mailgun (Mailgun is only required for the weekly digest email feature)
2. Install dependencies
npm install
3. Set up Supabase
Create a project at supabase.com, then run every
migration in supabase/migrations/ in filename order against it
(via the Supabase SQL editor, or the Supabase CLI if you have one set
up):
0001_init.sql
0002_features.sql
0003_browserbase.sql
0004_phone_numbers.sql
0005_benefits_state.sql
0006_saved_specialists.sql
0007_direct_messages.sql
0010_specialist_provenance.sql
0011_expert_call_requests.sql
0012_inapp_call.sql
0013_annotation_layer.sql
0014_blind_eval.sql
0015_claim_timeout.sql
0016_remove_digest.sql
(Note the gap between 0007 and 0010 — there is no 0008 or 0009; that's expected, not a missing file.)
From your Supabase project's API settings, grab the project URL, anon key, and service role key.
4. Get the rest of your credentials
- Anthropic — create an API key at console.anthropic.com.
- Browserbase — create a project and API key at browserbase.com. Used to power the live specialist and benefits search (billed through Browserbase, not a separate Anthropic key).
- Daily.co — create an API key at dashboard.daily.co. No other dashboard setup needed — rooms and tokens are created per-request by the app.
- Vapi — create an account at vapi.ai and get
your private API key (
VAPI_API_KEY) and public key (NEXT_PUBLIC_VAPI_PUBLIC_KEY). - Terac — get your API key and project ID from your Terac hackathon dashboard.
- Mailgun — only needed if you want the weekly digest emails to actually send; otherwise leave the placeholder values.
5. Configure environment variables
Copy .env.example to .env.local and fill in every value:
cp .env.example .env.local
All variables Compass actually reads are documented inline in
.env.example (Supabase, Anthropic, Mailgun/digest, Terac, Daily.co,
Vapi, Browserbase).
6. One-time setup script: create the Vapi assistant
The comfort-companion assistant has to be minted once via the Vapi API
before NEXT_PUBLIC_VAPI_ASSISTANT_ID exists:
VAPI_API_KEY=your-key node scripts/create-vapi-assistant.mjs
This prints an assistant ID — paste it into .env.local as
NEXT_PUBLIC_VAPI_ASSISTANT_ID. If you need to change the assistant's
behavior later without recreating it, use
node scripts/update-vapi-assistant.mjs instead (it patches in place and
wires up serverUrl so Vapi's end-of-call webhook reaches
/api/vapi/webhook).
7. Run the app
npm run dev
Visit http://localhost:3000.
8. (Optional) Run the before/after eval pipeline
Requires real annotation data already in your annotations table (i.e.
real live-connect calls have happened and at least one annotator has
submitted a correction) and a real ANTHROPIC_API_KEY / Supabase
service role key in .env.local:
npx tsx scripts/build-improved-prompt.mjs # writes scripts/generated/prompt-versions.json
npx tsx scripts/run-eval.mjs # writes scripts/generated/eval-pairs.json
npx tsx scripts/collect-blind-eval.mjs # publishes a blind Terac opportunity
npx tsx scripts/report-results.mjs # prints the v1 vs v2 win rate
Each script must be run in order — each one consumes the previous
script's output file. collect-blind-eval.mjs spends real Terac
opportunity credits.
Known Limitations
- General-population annotators only. Per the Terac challenge's design, there's no clinical screening on who reviews a parent's call summary — just an optional free-text field. This is a constraint of the hackathon's annotator pool, not a production safety decision; a real deployment would need a much more deliberate consent and reviewer-vetting design before showing anyone details from an emotionally difficult conversation.
- The before/after eval pipeline has been run once, on 8 scenarios. The 100% win rate reflects a single 8-comparison run, not a large-scale study — a real production rollout would want a much bigger held-out set before trusting that number generally.
- Live-connect has several points of external dependency (Daily.co,
Vapi, Terac, all chained together) — if any one of Daily room
creation, the Vapi webhook delivery, or Terac's opportunity
creation/launch fails, the request is marked
failedand surfaced to the parent, but the feature is only as reliable as the slowest/least available of those three services on a given day. - Terac launches are hard-capped at $20 per opportunity as a safety
rail in
src/lib/terac/client.ts— if real-world pricing ever exceeds that, the launch is refused rather than silently overspending, and the request is markedfailed. - Poke's actual text-handling logic is outside this repo. Compass only collects and stores the parent's phone number; the conversational behavior lives in Poke's hosted recipe.
- IEP Coach and most directory/benefit lookups have no persistence requirement by design — they're meant to be quick, stateless tools, not records systems.
Team / Branch Structure
The Poke text-assistant feature (phone number collection in onboarding
and settings) was built independently by a teammate and merged in
alongside the Terac/live-connect work, which was developed on its own
feature branch in parallel. Both landed on main together; the
migration numbering reflects that interleaving (e.g. 0004_phone_numbers
sits between earlier general-feature migrations and the later
Terac-specific ones added once the live-connect branch merged).
Analysis
View
Metric
- 30
- 19
- 6
Figures cover GitHub contributors during the hackathon window. A co-authored commit counts in full for each author, so per-member totals add up to more than the whole-team figures.
Technology
- AnthropicIn code
- CSSIn code
- HTMLIn code
- JavaScriptIn code
- Next.jsIn code
- ReactIn code
- SQLIn code
- SupabaseIn code
- Tailwind CSSIn code
- TypeScriptIn code
10 of 10 appear in the indexed code.
AI coding agents
- Claude CodeConfig · Commits
Detected from committed agent config files and commit authorship. Absence of a signal is not proof an agent was unused.
Codebase size
Source size
1.2 MB
Source files
164
Counts recognized source files only; vendored directories, binaries and lockfiles are excluded, so this is smaller than the repository on disk.
Repository
recalculator/Compass
206 files · 19.0 MB · @ 0563013
Structure
Interface
63 files · 31%Screens, components and styles rendered to the user.
API & routing
26 files · 13%Request entry points: routes, handlers and controllers.
Application logic
35 files · 17%Domain rules, services and shared utilities.
Data & schema
18 files · 9%Schema definitions, migrations and data access.
Supporting
Layers are inferred from where files sit in the tree, not from reading the code. A project that names its directories unconventionally will read oddly here — open the file browser to check anything the diagram implies.
Languages
- Markdown69%
- TypeScript22%
- HTML6%
- SQL2%
- CSS0%
- JavaScript0%
Share of indexed source by file size. Binary and vendored files are excluded.
Dependencies
package.json
npm · 24- @anthropic-ai/sdk
- @browserbasehq/stagehand
- @daily-co/daily-js
- @supabase/ssr
- @supabase/supabase-js
- @vapi-ai/web
- clsx
- date-fns
- lucide-react
- next
- pdfkit
- react
- react-dom
- zod
- +10 more
Declared in the repository’s manifests at the indexed commit. A declared package is not proof it is used, and runtime dependencies are listed first.
This project’s features have not been analysed yet.
Export this project's context (description, README, evidence, key source files) to chat with an AI agent elsewhere.