Thrive PM · Product Note

Kickoff & Sourcing Agents — Learnings

Kickoff & Sourcing Agents — Learnings

Last updated: 2026-08-05
Inputs covered: feedback through 2026-08-11 (9 users; deep-dive calls: Jim Rosen + Maria Grace 8/4, Matt Davis 8/6, Joe Ganim 8/7, Alex Reisley 8/11); usage/cost data through 2026-07-17 (24 pilot kickoffs; internal test runs excluded); outcomes data on 9 projects (2026-08-05); business case as revised 2026-07-30
Raw log: ../feedback/kickoff-sourcing-agents.md
Briefs & business case: ../artifacts/ (kickoff + sourcing briefs 2026-05-18; business case + ROI 2026-07-01, revised 2026-07-30)

1. Stakeholder summary

Where we are. 12 True Search pilot users have run 24 kickoffs since 6/22, generating ~2,700 candidates across 18 completed sourcing runs (avg ~152 per search). Sentiment is strongly positive on the core concept: "I just used the Kickoff agent and it was great!!" (Jim Rosen), "i am very impressed by the four outputs" (Matt Davis). No user has rejected the workflow; all feedback is bugs, polish, or "make it do more."

Against the brief's use cases. The sourcing agent is delivering real pipeline: Matt Davis added 23 candidates to a live search in ~1 hour — meeting the sourcing brief's "<1 hour to an initial market map" target on its first outing (GTM baseline: ~16 hours); Joe Ganim says it surfaces "bullseye profiles + some I missed" even entering mid-project. 19 of 24 pilot kickoffs proceeded to a sourcing run (~79%), tracking above the brief's 70% attach target. Reported hit rates are 10–25% of generated profiles (Tim McDonald), with precision weakest on harder or loosely-specified specs (Steve Tutelman: adjacent titles like "VP revenue accounting" for a VP Revenue search) — the brief's quality bar (>20% of added candidates reach Pursuing or above within 60 days) is now measured on nine projects — and met: 261 agent-generated candidates sit in real pipelines (19% of the 1,357 candidates in the latest captured runs), 88 have moved past Research, and 54 — 20.7% of adds — have reached Pursuing or above. That clears the brief's >20% bar, and it's still a floor: these searches remain open and the 60-day window hasn't elapsed. Vinyl Equity (50%), Zarminali (45%), and Maria's Gig Safe VP Eng search (42%) lead on adds reaching Pursuing+. Usage styles differ sharply (Steve bulk-added ~100% of a run; Joe curated 4%), and on every project the agent's 1–5 score predicts which candidates recruiters add and advance (added candidates average 3.1–3.8 vs 2.0–2.5 for the rest). The kickoff documents get consistent good first impressions, but no user has yet reviewed them in depth — document quality is unvalidated.

Against the business case. Inference cost is running as modeled: $18.69 total across the pilot, mean ~$0.78/run — consistent with the ~$7.8K/yr inference assumption in the business case (2026-07-01, revised 2026-07-30: base case ~$603K/yr gross / ~$596K net recovered capacity vs ~$120K build, ~2.4-month payback, ~396% Y1 ROI). The value side rests on two unvalidated inputs — the 60% time-savings rate and the 25% heavy-touch practice mix; Matt's 23-candidates-in-an-hour is the first supporting datapoint on time saved. The planned 30-day user survey (time saved by practice) and a Domo practice-mix pull are the validation path, with Pendo to follow.

From Alex Reisley's 8/11 deep-dive call: the fifth deep-dive confirms the scoring pattern with two more live examples — a Palantir-lifer CTO matched at five stars ("way too big for this role… not a good match at all") and an ex-CEO now CTO at C3 rated a strong fit — and, usefully, puts numbers on the tenure signal: >10 years at one company reads as a red flag, <1 year as dubious unless the prior track record is strong. Two new asks with weight behind them: boolean/expanded title matching ("that's how I usually do my LinkedIn sourcing… a lot of candidates don't have these exact titles") and past-project sourcing ("I use it for everything basically — a top-three parity feature for Thrive 2.0"). He also hit an unresolved failure: a niche multi-location search that errors and returns nothing (suspected strict-location empty state — under investigation). Enthusiastic on calibration candidates, including annotating why someone is a good example.

From Joe Ganim's 8/7 deep-dive call (ran the agent live): the clearest articulation yet of what the agent is for — "the first names that pull up are the most obvious names, which is what you want… the quicker we can get those into the project, the more time it saves. Finding the diamonds in the rough is really hard for any automated workflow — that's where our job is." That reframes the value proposition: speed-to-the-obvious-names is the product; diamonds-in-the-rough is the human's edge (and the stretch goal for notes ingestion + enrichment). He called update-call notes "probably the most lucrative" future source, quantified location (irrelevant on ~70% of searches, decisive on 10–15%), and added the key scoring constraint that seniority-vs-size calibration is per-project — his CPO client wants "miss high, not low," the opposite of Jim's search. Sentiment: "genuinely, it is a completely different product than it was [weeks] ago."

From Matt Davis's 8/6 deep-dive call: a moderator for when the agent shines — title specificity. His customer-success search "was actually really good… it knew that I needed big companies… it really got the gist right," because "customer success generally means the same [thing]" while "VP sales could mean a million different things" — for ambiguous-title functions, "you're going to need to get into the notes." He also confirmed the already-in-project bug from the user side, and asked for pipeline-stage selection on add, Metaview call-notes ingestion with automatic re-runs, and real-time kickoff-doc editing — all now on the roadmap (T5-16, T4-17, T5-11).

From the 8/4 deep-dive calls (Jim Rosen, Maria Grace). The precision problem now has a diagnosed root cause: the agent matches title×company literally, without calibrating seniority to company size or weighting current vs past roles — so listing DaVita as a target company surfaces DaVita's actual COO when the search wants former DaVita division VPs now running small PE-backed companies. Both users independently validated the market-map use case as-is ("if you could just have a market map created to start the search… that would save so much time off the bat" — Jim; "100%… we've been trying to show them this is really the only talent in Texas" — Maria), and both want an optional feedback loop where not-a-fit reasons and client-call notes recalibrate the next generation. Sentiment stays strong: "this is amazing already" (Jim).

Top risks. (1) A replicated bug where edited search criteria revert/reset — two independent reports — undermines trust in the core iteration loop. (2) Chrome's pop-up blocker silently swallows the kickoff output window; usage data suggests up to 3 of 12 pilot users hit this, and 2 of them never returned. Both are fixable and neither challenges the concept.

Next steps. Ticket the criteria-persistence bug and pop-up handling; collect the deferred document-quality feedback (4 users owe it); define hit-rate and time-saved measurement before scaling the pilot.

Theme A — Search criteria edits don't persist (bug) — HIGH

  • Evidence: Samir Sidhu 7/13 ("it keeps reverting back" — replicated live on a call); Maria Grace 7/24 (added companies are dropped back to the originally-generated list when she edits another field like location).
  • Type: Bug. Priority: high — two independent reports in two weeks, replicated, and it breaks an explicit brief requirement ("Changes are saved in the agent so users can revisit the criteria page, make edits, and re-run") — the edit-criteria-and-rerun loop users (Joe, Steve) explicitly want to use.
  • Action: File an engineering ticket with both repro stories; Maria's version (edit field B, lose additions to field A) is the more specific repro path and suggests the form rehydrates from the generated snapshot instead of the user's edits.

Theme B — Pop-up blocker silently kills the kickoff output — HIGH

  • Evidence: Jim Rosen 7/17 (hit kickoff, "didn't see anything happen"; diagnosed Chrome's pop-up blocker himself days later). Usage data corroborates: his 7/14 Zarminali run shows 0 candidates/no status, then success on retry 7/17. Mark Sylvester (6/29) and Jennifer Brannigan (7/6) each have exactly one 0-candidate/no-status run and no activity since — plausibly the same failure, experienced as "the product did nothing."
  • Type: Bug/UX (activation-killer). Priority: high — it can end a pilot user's journey on day one, and it's invisible in our success metrics without this cross-check.
  • Action: Detect the blocked window and show an inline fallback link (or stop opening a new window at all). Interim: add a pop-up-blocker note to pilot onboarding. Also confirm with Mark and Jennifer whether this is what they hit (see §4).

Theme C — Review workflow: dismiss, capture why, recalibrate — MEDIUM→HIGH (design in progress)

  • Evidence: Matt Davis 6/23 thread — wants a "not-a-fit"/hide action like LinkedIn Recruiter; returning to the agent forces re-wading through already-reviewed profiles; rejections must stay retrievable "in case the spec shifts later on." Extended by the 8/4 calls: Maria wants yes/no/maybe with an optional reason ("if it's you have the option to, instead of you have to"); Jim wants to dump client-call notes in for re-generation, with the ability to discount specific client comments ("clients will make up logical reasons for why we don't like something that's just a gut feel").
  • Type: UX/feature — now a feedback-loop feature, not just visual hygiene. Priority: raised to medium-high — three users have independently described the same loop (dismiss → reason → recalibrate), and it's the mechanism by which the agent learns the nuance in Themes D/E. Matt 8/6 added the zero-effort version: sync Metaview call notes so the agent re-runs on the latest client feedback automatically ("every client call we get feedback… what worked, what didn't work about candidates") — roadmap T4-17. Joe 8/7 joined (4 users on the loop) and extended the dismiss action with save-for-later (parity with the Thrive recommendations tab) plus reason-feeding ("I'm going to mark him down because of history — they don't want people from large companies"); Metaview caveats from him: he uses it for client calls but not candidate calls, and finds it clunky/slow.
  • Action: Continue the design exploration with the loop framing: persistent reviewed/rejected state, optional reason capture, and reasons + call notes feeding re-runs.

Theme D — Sourcing breadth: go beyond the listed target companies — MEDIUM

  • Evidence: Jim Rosen 7/17 (candidates come almost entirely from the entered target companies; wants expansion to similar companies, and category-level targets like "PE-backed dental businesses" as target companies/industries); Mariyam Fatima 6/29 (broad keywords alongside titles — generic title + domain keywords like lending/capital/banking). Now measured: 57–72% of generated candidates come from listed target companies on every one of the eight analyzed projects (pilot-outcomes analysis, 2026-08-05) — Jim's observation quantified and general, not healthcare-specific. Note: Thrive projects already carry benchmark candidates and benchmark jobs (Zarminali has 3 + 2) that the agent doesn't consume — an existing data source for Maria's calibration-profiles ask and her similar-past-searches idea.
  • Type: Feature. Priority: medium — this is the most-requested capability expansion and directly increases candidate yield per search, but note the tension with Theme E (precision) — expansion needs a relevance guardrail.
  • Alex Reisley 8/11 joined the similar-company ask ("intelligence to grab candidates who might be at other startupy companies") and named the blocking data gap precisely: Thrive doesn't associate asset class / revenue range with candidates, only (sometimes) via manual tags — he wants automatic tagging ("with no effort at all… it just automatically tags that candidate"), which is the enrichment enabler plus the parked auto-tagging item, now a 2-user ask. He also brought a new title-matching ask: boolean/expanded titles ("senior director, enterprise AI engineering is not that common… it's usually VP engineering, AI"), and made past-project sourcing a heavy-usage data point ("I use it for everything").
  • Joe Ganim 8/7 joined the similar-company ask ("what are similar adjacent companies… would be super helpful") and named update-call notes "probably the most lucrative" future source — with the practical caveat that it depends on recruiters diligently uploading notes. On PE-portfolio carveouts he's a counterweight: "fantastic to have, just not mission critical" (his workaround: ask Claude for the portfolio list and pin it to the criteria).
  • 8/4 calls sharpened the ask into three tiers: (1) example-based expansion — give one exemplar company ("Forefront Dermatology") and have the agent find its competitors, possibly via external data like PitchBook (Jim); (2) reasoning search — express the actual spec in natural language ("former DaVita division VPs now COO/SVP ops at PE-backed healthcare companies") and let the agent plan the query (Jim); (3) similar past searches — surface comparable Thrive searches by company stage/revenue/PE-backing to mine (Maria). Maria also asked for PE-backer identification on candidates' companies to avoid off-limits conflicts (Vista/Bain).
  • Action: Scope "similar companies" expansion and industry/category targets as a candidate roadmap item; the third-party data integration already underway is the enabler for tiers 1–2. Demand is confirmed general, not healthcare-specific.

Theme E — Precision: too many out-of-bounds profiles — HIGH (root cause diagnosed)

  • Evidence: Steve Tutelman 7/16 (adjacent-function titles; "probably pushed too many profiles that were outside the bounds"); Tim McDonald 7/13 (25% hit rate on OpenX CEO, ~10% on the harder Triton Chair spec; outcomes data measured 21% and 3%). Root cause diagnosed on the 8/4 calls: literal title×company matching with no calibration for (a) seniority relative to company size — DaVita's actual COO surfaces when the spec wants people a level or two down at giant companies who'd step up at a small one (Jim; Maria on Google CTO vs small-company CTO: "we'd want to scale it down by however many levels — absolutely, 100%"); (b) current vs past role weighting — past-CTO-now-advisor ranked as a match (Maria); (c) tenure/flight-risk signals — 20-year Concentra lifers "probably not leaving ever" (Jim); (d) hard location exclusion where a soft preference/downrank is wanted (Maria).
  • Thresholds (Alex Reisley 8/11): third and fourth perfect-score misses observed live (Palantir-lifer CTO at 5 stars; ex-CEO C3 CTO rated strong) — same root cause as Zarminali. He supplied concrete starting values for the tenure signal: >10 years single-company ≈ red flag; <1 year in role ≈ dubious (with a strong-prior-decade exception). He'd also downgrade rather than exclude — consistent with soft-preference scoring everywhere.
  • Constraint (Joe Ganim 8/7): the seniority-vs-size correction has no global direction — "a VP at Meta is very different than a VP at [a small company]," but some clients explicitly don't want hyperscaler VPs ("you can hide in a big company"), and Joe's CPO client prefers missing high over missing low while Jim's wanted step-up candidates. The calibration signal is stated by clients on calls, so scoring v2 ultimately needs per-project inputs (notes, calibration profiles) rather than one rule. Also quantified: location is irrelevant on ~70% of searches and decisive on 10–15% — confirming must-have/nice-to-have as the right control shape.
  • Moderator (Matt Davis 8/6): precision tracks title specificity. Specific-function titles (chief customer officer) produce "a lot of relevant results" with today's matching; ambiguous titles (VP Sales, marketing) will stay weak until the agent reads notes/scorecards (Theme D tier 2 / roadmap T3-10). This predicts which searches hit and suggests segmenting quality metrics by title ambiguity.
  • Type: Quality/scoring. Priority: raised to high — it's the top driver of wasted review time, three users have converged on it, and the fixes are specifiable (scoring inputs, not vague "be smarter").
  • Hard evidence (2026-08-05): the Zarminali capture turned Jim's call into a labeled example set (pilot-analysis/2026-08-05_zarminali-labeled-examples.md). All three of his "not a fit" candidates (sitting COOs of DaVita/Concentra/LifeStance) scored a perfect 5; his one "on-target" example (VP Ops at Envision, who really did reach Pursuing) also scored 5 — the score can't tell them apart. Sharpest finding: the agent generated the stated bullseye pattern (DaVita Division VPs) but scored it 4, below the giant-co COOs — a scoring inversion. And the criteria agent had captured the nuance (the "Division Vice President" title, the "$150M–$500M inflection" signal), so the failure lives in scoring/weighting, not criteria extraction. Doc includes 3 regression assertions for scoring v2.
  • Action: (1) Quick win — Jim's interim ask: add input guidance telling users the agent only searches the listed companies+titles, so "be exhaustive" (prompt/UI copy, no model work). (2) Scope scoring v2: seniority-vs-company-size scaling, current-role weighting, tenure signal, location as soft preference — use the labeled examples as the acceptance tests. (3) Validate against the outcomes data — the score already predicts adds/progression, so calibration improvements are measurable with the pilot-outcomes pipeline.

Theme F — Trust signals & data quality in results — INVESTIGATE

  • Evidence: Mariyam Fatima 6/29 — (a) how are "Unknown"-location profiles filtered, given Thrive often misses a location that's on LinkedIn? (b) surface whether a candidate is active in another project/OL or reached CI in the last 18 months — a "Best Match + recently in touch with True" section for quick intros.
  • Type: Data-quality + feature. Priority: investigate — single reporter so far, but the recency/relationship idea aligns with the TRM's core differentiator (warm networks) and could be high-leverage.
  • Action: Answer the Unknown-location filtering question factually (check with eng how the filter treats nulls — silent exclusion would shrink candidate pools). Float the "recently in touch" section with 2–3 other users.

Theme G — Kickoff should accept more input types — LOW

  • Evidence: Mariyam Fatima 6/29 — clients send JDs and notes ahead of kickoff; want to add multiple docs alongside kickoff notes.
  • Type: Feature. Priority: low/medium — single report, but low-friction inputs compound adoption.
  • Action: Check current input handling; if multiple docs already work, this is a discoverability fix, not a feature.

Theme H — Document quality is unvalidated — FOLLOW-UP, not a defect

  • Evidence: Joe, Tim, Alex, and Matt all gave positive first-glance reactions but explicitly deferred deep review of the generated documents. Nobody has reported on document accuracy/completeness yet.
  • Type: Evidence gap. Priority: high as a follow-up — kickoff documents are half the feature and half the business-case time savings.
  • Action: Chase the four owed reviews (see §4); consider a lightweight rubric (accuracy / completeness / edits needed) so responses are comparable. Partial signal 8/4: Maria — docs are "definitely a starting point… I'll just edit it, incorporate it… already been really helpful."

Theme I — Candidate card lacks decision-critical context — MEDIUM (new 8/4)

  • Evidence: Maria Grace call — she opens every candidate's LinkedIn in a tab to check tenure ("somebody's been there for three months… I wish I didn't even have to click into that"); wants a LinkedIn-Recruiter-style tenure dropdown ("with this company from this year to present") and an experience breakdown on the card. The existing per-candidate summary was explicitly helpful. Related: Mariyam's 6/29 "Best Match + recently in touch" ask (Theme F) is the same pattern — surface the deciding signal on the card.
  • Type: UX. Priority: medium — directly cuts per-candidate review time, which is where sourcing time savings accrue.
  • Action: Add tenure/date ranges and a compact employment timeline to the result card. Employment dates are already in the sourcing payload, so this is largely display work.

Theme J — Write-back to Thrive: target companies yes, docs manual — MEDIUM (new 8/4)

  • Evidence: Maria Grace call — auto-pushed docs "check the audit stuff off the box, but I still have to go in and edit it anyway"; prefers dropping docs in "if and when ready." Pushing target companies to the Thrive project would be "very helpful for sure" — she re-enters them manually today.
  • Extended by Matt Davis 8/6: the add action itself needs to land where the work happens — he wants to choose the pipeline stage when adding (including straight to Rejected with an off-limits reason, so the review is recorded and the person resurfaces when their status changes). He also asked what the Thrive strategy tab feeds into ("if I have really thoughtful strategy information in there, does that affect the recommendations?") — the write-back should be a two-way loop, not a one-way push. And a data-quality dependency: tags should be auto-generated with a human in the loop (parked with core product).
  • Type: Feature/workflow. Priority: medium — target-company write-back is a clear small win; the docs-push change needs care (auto-push currently satisfies audit/data-completeness requirements).
  • Action: Scope target-company write-back (user-controlled toggle, per Eleyni's on-call framing) and stage-selection on add (roadmap T5-16). Validate the docs-push change with 2–3 more users + whoever owns the audit requirement before changing the default.

Validated so far (keep doing): the market-map-at-kickoff use case (Jim + Maria, 8/4 — independently and emphatically); criteria generation ("I wouldn't even think to type all that in and it's already there" — Maria); per-candidate summaries; docs as an editable starting point.

3. Usage & outcome data

From 2026-07-17_sourcing-agent-llm-cost.csv — 24 pilot kickoffs, 6/22–7/17 (the export also contains 2 internal Thrive test runs, excluded from all metrics here):

Metric Value Interpretation
Pilot users / kickoffs 12 users, 24 kickoffs Repeat usage from 6 users; Joe Ganim heaviest (5). Real adoption, not one-and-done demos.
Completed sourcing runs (pilot) 18 of 24 1 failed (Joe/MRI CISO, 6/22), 5 with no run status and 0 candidates. 19/24 kickoffs reached a sourcing run (~79% attach) — early signal above the brief's 70% target and the business case's 75% input, small sample.
Candidates generated (pilot) 2,731 (avg ~152/completed run) Wide range: 12–374. Low-count runs (12–26) worth a look — thin spec or thin market?
LLM cost $18.69 pilot total; mean ~$0.78/run, max $3.34 Matches the business case's ~$7.8K/yr inference assumption (built from this same data). Cost is a non-issue.
Zero-result, no-status runs 5 pilot runs across 3 users Jim's two Zarminali attempts are confirmed pop-up blocker. His later Summit Partners run (7/17, after he'd diagnosed the blocker) plus Mark Sylvester's and Jennifer Brannigan's (1 each, never returned) are unexplained — churn risk (Theme B).
Reported hit rate 10–25% (Tim, 2 searches); 23/230 = 10% added (Matt) Self-reported. Usage patterns differ sharply: Steve bulk-added ~100% of a run; Matt curated ~10%. "Hit rate" needs the stage-progression lens, not just adds.
Measured outcomes (9 projects) 261 of 1,357 latest-run candidates added (19%); 88 moved past Research (34% of adds); 54 Pursuing+ (20.7% of adds — brief's >20% bar met, as a floor); $0.04 LLM cost per added candidate Pilot-outcomes pipeline, 2026-08-05. Tim's two searches corroborate his self-reports directionally (OpenX 21% added vs "~25%", harder-spec Triton 3% vs "~10%"). Curated usage converts best: Vinyl 50%, Zarminali 45%, Gig Safe 42% of adds at Pursuing+. Score calibration positive on every project with adds. Zarminali also yielded a labeled example set proving the scoring-inversion root cause (pilot-analysis/2026-08-05_zarminali-labeled-examples.md). Caveat: sourcing captures hold the latest run only (3 of 9 jobs had earlier runs). Full doc in pilot-analysis/2026-08-05_pilot-outcomes.md.

Collection gaps: Pendo not live yet (funnel, session behavior). No candidates-added-to-project / outreach / stage-progression data — so the sourcing brief's quality metric (>20% of added candidates advance past Identified within 60 days) can't be computed yet. No time-saved measurement (the business case's 60% time-savings rate and 25% heavy-touch practice mix remain unvalidated — Domo practice distribution + 30-day survey are the planned instruments).

4. Open questions & follow-ups

  • [ ] Eng: ticket + investigate criteria-persistence bug (Theme A) — use Maria's repro (add companies → edit location → additions lost).
  • [ ] Eng (from 8/11 call): investigate Alex Reisley's erroring niche search — errors instead of returning results, retried live twice; suspected strict multi-location filter producing an empty state that surfaces as an error (Themes B/E adjacent; Eleyni promised to look into it).
  • [ ] Eng/design: ticket pop-up-blocker handling (Theme B); interim onboarding note.
  • [ ] Mark Sylvester & Jennifer Brannigan: one 0-result run each, then silence — ask what happened (suspect pop-up blocker). Recovers 2 of 12 pilot users if so.
  • [ ] Joe, Tim, Alex, Matt: collect the promised document-quality reviews; Alex asked for a call. Consider a short rubric.
  • [ ] Tim / Steve: get 3–5 examples of "wrong" profiles to characterize the precision misses (Theme E). Jim's 8/4 walkthrough delivered this for his search; Tim/Steve examples would confirm the same root cause outside healthcare.
  • [ ] Eng/product (from 8/4 calls): scope scoring v2 — seniority-vs-company-size scaling, current-role weighting, tenure signal, location as soft preference (Theme E).
  • [ ] Quick win (from 8/4 calls): draft the criteria-screen guidance copy — "the agent searches only the companies and titles you list; be exhaustive" (Theme E interim fix).
  • [ ] Validate (from 8/4 calls): docs auto-push default change + target-company write-back with 2–3 more users and the audit-requirement owner (Theme J).
  • [ ] Steve: after his second search, compare experience vs first (learning-curve signal; input-guidance implications).
  • [ ] Joe: did editing criteria to match the pivoted PSG spec work? (His PSG run generated only 12 candidates.) Also overlaps Theme A.
  • [ ] Mariyam's location question: get the factual answer on how "Unknown" locations are filtered; respond to her.
  • [ ] Jessica Hirst & Abby Mitchell: ran kickoffs but no feedback yet — solicit.
  • [ ] Data: collect hit rate and stage progression per project (the brief's >20%-advance metric) — in progress: the pilot-outcomes pipeline (analyze-pilot-outcomes skill) has measured 9 projects as of 2026-08-05 (RTA, Rivet, Vinyl, OpenX, TA, Triton, Turn/River, Zarminali, Gig Safe); 10 completed pilot searches remain to capture (top of list: Sectigo, Stout, PlanHub, Ellipsis, PSG).
  • [ ] Data: Pendo instrumentation live → funnel from kickoff → sourcing run → candidate added.
  • [ ] Data: time-saved measurement (30-day survey) + Domo practice-mix pull to validate the two load-bearing business-case assumptions (60% time-savings rate, 25% heavy-touch mix).
  • [ ] Roadmap validation: float "similar companies / industry targets" (Theme D) and "recently in touch with True" (Theme F) with 2–3 more users each.

5. PM decisions & notes

(Space for Eleyni — preserved verbatim across regenerations.)

6. Change log

  • 2026-08-11: Synthesized Alex Reisley's 8/11 call (transcript in data dir). Theme E gained two more live perfect-score misses + concrete tenure thresholds (>10y red flag, <1y dubious); Theme D gained boolean/expanded titles (new ask), his similar-company + past-projects-heavy-usage evidence, and the candidate-level asset-class/revenue tagging gap. New follow-up: investigate his erroring niche search. Roadmap deliberately NOT updated — recommendations delivered to Eleyni for a decision (per her instruction), not applied.
  • 2026-08-10: Synthesized Joe Ganim's 8/7 call (ran the agent live on the rescoped Stout CPO search; transcript in data dir). §1 gained his value-prop framing (speed-to-obvious-names is the product; diamonds-in-the-rough is the human's edge). Theme E gained the per-project calibration constraint + location quantification (irrelevant ~70% / decisive 10–15%); Theme C gained save-for-later + reason-feeding (4 users on the loop); Theme D gained his adjacent-companies + notes-most-lucrative evidence and a carveouts counterweight. Roadmap regenerated as 2026-08-07r2 (75 items): T4-01 extended to not-a-fit and save-for-later; multi-user evidence on T3-08/T3-10/T2-02/T2-05/T4-17/T3-12.
  • 2026-08-07: Synthesized Matt Davis's 8/6 call (transcript in data dir). Title-specificity moderator added to Theme E (specific titles work today; ambiguous ones need notes); Metaview auto-ingestion added to Theme C; stage-selection-on-add + strategy-tab loop + auto-tagging added to Theme J. Roadmap regenerated (2026-08-07 deck + xlsx, 75 items): new T5-16 (stage on add) and T4-17 (Metaview); T5-11 (real-time doc editing) promoted to Next/P1; evidence strengthened on T1-03/T3-02/T3-10; PL-05 auto-tagging parked with core product.
  • 2026-08-05 (later): Zarminali + Gig Safe (Maria's VP Eng search) captured → 9 projects measured; aggregate Pursuing+ hit 20.7% of adds — the brief's >20% bar is met (as a floor). Built the Zarminali labeled example set from Jim's call: his three "not a fit" candidates all scored 5 while the generated DaVita Division-VP profiles (his stated bullseye) scored 4 — scoring inversion confirmed; criteria extraction had captured the nuance, so the failure is isolated to scoring/weighting. Also surfaced: Thrive benchmark candidates/jobs exist on projects and are unconsumed by the agent.
  • 2026-08-05: Synthesized the 8/4 deep-dive calls (Jim Rosen, Maria Grace — full transcripts in the data dir). Theme E raised to HIGH with a diagnosed root cause (no seniority-vs-company-size calibration, current-role weighting, tenure signal; hard location filter); Theme C reframed as a recalibration feedback loop and raised; Theme D sharpened into three tiers (example-based expansion / reasoning search / similar past searches) + PE-backer awareness. New Themes I (candidate-card tenure/context) and J (write-back: target companies yes, docs manual). Market map + criteria generation validated as use cases. 4 follow-ups added, incl. two quick wins.
  • 2026-07-31 (later): Added Triton Chair and Turn/River CCO → 7 projects measured (211/1,095 added, 19%; 32 Pursuing+, 15% of adds). Tim's easy-vs-hard-spec self-report confirmed directionally (OpenX 21% vs Triton 3%); Matt's curate-then-advance pattern repeats; Triton shows the sharpest score calibration yet (added avg 4.67 vs 2.28).
  • 2026-07-31: Outcomes analysis expanded to 5 projects (added Rivet CTO, Vinyl Head of Mktg, OpenX CEO, TA Associates). Aggregate: 25% of latest-run candidates added, 15% of adds at Pursuing+ (floor; searches open). Target-company dependence confirmed general (57–71% everywhere); score calibration positive on every project; OpenX's measured 21% corroborates Tim's self-report. New anomaly: Samir's TA run (latest of many re-runs — he's the Theme A criteria-bug reporter) generated only 17 candidates, 0 added.
  • 2026-07-30: Initial synthesis. Inputs: 9 users' feedback (6/23–7/24) + LLM cost/usage CSV through 7/17. Opened 13 follow-ups. Themes A–H established.
  • 2026-07-30 (outcomes pipeline): Built analyze_pilot_outcomes.py + analyze-pilot-outcomes skill — per-project agent→pipeline analysis from captured network responses (no Thrive API). First project measured (Ron Turley CRO): 99% of the latest run added (bulk-add pattern), 6% Pursuing+ so far, 71% of generated candidates from listed target companies (quantifies Theme D), score calibration positive. §3 gained a measured-outcomes row; hit-rate follow-up now in progress.
  • 2026-07-30 (later still): Excluded internal test runs (Eleyni's 2 Thrive test kickoffs) from all usage metrics per PM direction — headline counts now 24 pilot kickoffs from 6/22; rule encoded in the synthesize-learnings skill for future runs.
  • 2026-07-30 (later): Re-synthesized against the relocated artifacts (eleyni-product/artifacts/) and the revised business case (xlsx updated: 60% time-savings rate, 75% sourcing attach → base $603K gross / $596K net, ~2.4-mo payback, ~396% Y1 ROI; conservative floor now Y1-positive). Grounded §1 in the briefs' success metrics (70% attach, <1hr market map, >20% advance); noted pilot attach ~79%; Theme A elevated as a brief-requirement violation. No new user feedback; no theme or follow-up changes otherwise.