Copy-on-Write Scoring: Application-Specific Agent Evaluations
Trustworthy deployment of LLM-based agents in software systems requires evaluating how they perform on application-specific workflows, with enough granularity to localize where they succeed and fail. Yet existing agent evaluation mechanisms
https://arxiv.org/abs/2607.14336v1 ↗Thesis fit
Good fit
Within your typical scope; diligence still required.
In your usual scope
Idea match
Light
How close the company’s idea is to your thesis statement
Sector
agents, llm, ai
Overlap with sectors you care about
Geography
Unknown
Location unknown — scores 0
Your thesis: “We back exceptional technical founders building AI-first products and infrastructure, deploying $100K checks within 24 hours.”
Founder ↑ improving
Traction → stable
Idea vs market ↓ declining
▸ Add / edit details
Correct facts used on the next screening or memo.
People
J
Joanna Roy
1.0
low confidence
S
Sven Hoelzel
2.0
low confidence
J
J. Roy
7.0
low confidence
Activity & evidence
Similar baseline plays (YC · idea space)
Arga Labs
· Active
Real-world sandboxes to test agents and agent-facing software
founders not scraped yet
Gumloop
· Active
A no-code platform for creating agents and automating workflows with…
founders not scraped yet
ProjectX
· Active
Agent native workspace for heavy parallel workflows on the web
founders not scraped yet