DeepSWE: Measuring Frontier Coding Agents on Original, Long-Horizon Engineering
DeepSWE is a benchmark of 113 original, long-horizon software engineering tasks for evaluating coding agents. Most public agentic coding benchmarks follow SWE-bench in mining merged fixes from public GitHub repositories, which creates two p
https://arxiv.org/abs/2607.07946v1 ↗Thesis fit
Good fit
Within your typical scope; diligence still required.
In your usual scope
Idea match
Moderate
How close the company’s idea is to your thesis statement
Sector
agents
Overlap with sectors you care about
Geography
Unknown
Location unknown — scores 0
Your thesis: “We back exceptional technical founders building AI-first products and infrastructure, deploying $100K checks within 24 hours.”
Founder ↑ improving
Traction → stable
Idea vs market → stable
▸ Add / edit details
Correct facts used on the next screening or memo.
People
W
Wenqi Huang
1.0
low confidence
C
Charley Lee
1.0
low confidence
L
Leonard Tng
1.0
low confidence
S
Serena Ge
1.0
low confidence