← back

DeepSWE: Measuring Frontier Coding Agents on Original, Long-Horizon Engineering

DeepSWE is a benchmark of 113 original, long-horizon software engineering tasks for evaluating coding agents. Most public agentic coding benchmarks follow SWE-bench in mining merged fixes from public GitHub repositories, which creates two p

https://arxiv.org/abs/2607.07946v1 ↗
Thesis fit
Good fit

Within your typical scope; diligence still required.

Edit thesis
In your usual scope
Idea match Moderate

How close the company’s idea is to your thesis statement

Sector agents

Overlap with sectors you care about

Geography Unknown

Location unknown — scores 0

Your thesis: “We back exceptional technical founders building AI-first products and infrastructure, deploying $100K checks within 24 hours.”

Founder ↑ improving Traction → stable Idea vs market → stable
Generate memo
Add / edit details

Correct facts used on the next screening or memo.

Similar baseline plays (YC · idea space)

Zenbu · Active
The extensible IDE for coding agents
founders not scraped yet
Magnitude · Active
The best coding agent for open models
founders not scraped yet
Continue · Acquired
Pioneering open-source coding agent
founders not scraped yet
sudocode · Active
orchestrate your coding agents with sudocode
founders not scraped yet
Compyle · Active
The coding agent that actually collaborates with you
founders not scraped yet