In-Place Tokenizer Expansion for Pre-trained LLMs
A tokenizer fixed at the start of pre-training allocates vocabulary in proportion to the pre-training corpus, reflecting the deployment priorities at that time. When those priorities shift, languages added later are split into many more tok
https://arxiv.org/abs/2607.15232v1 ↗Thesis fit
Good fit
Within your typical scope; diligence still required.
In your usual scope
Idea match
Light
How close the company’s idea is to your thesis statement
Sector
llm, ai
Overlap with sectors you care about
Geography
Outside
Outside your target regions — scores 0
Your thesis: “We back exceptional technical founders building AI-first products and infrastructure, deploying $100K checks within 24 hours.”
Founder → stable
Traction → stable
Idea vs market ↓ declining
▸ Add / edit details
Correct facts used on the next screening or memo.
People
J
Jimmy T. H. Smith
2.6
low confidence
T
Tarek Dakhran
2.2
low confidence
A
Alberto Cabrera
1.0
low confidence
S
Simon S. Lee
2.0
low confidence
P
Paul Pak
2.0
low confidence
A
Aditya Tadimeti
2.0
low confidence
T
Tim Seyde
2.0
low confidence
M
Maxime Labonne
2.0
low confidence
A
Alexander Amini
2.0
low confidence
M
Mathias Lechner
2.0
low confidence
A
A Cabrera
ES
1.0
low confidence
Activity & evidence
Similar baseline plays (YC · idea space)
Luel
· Active
Turning everyday words and actions into usable training data.
founders not scraped yet
Automorphic
· Active
Infuse knowledge into language models with just 10 samples
founders not scraped yet