Law is AI's next trillion dollar frontier.
Today's models can remember legal facts, but they lack judgment. They'll read a statute or summarize a case, but can't tell you whether a clause is too aggressive or a risk is worth taking. The senior partner's taste, style, and ability to blend facts with circumstances is exactly what you pay them for. And why it's so hard to measure and teach.
So today, we're launching Crosby Intelligence, to push the frontier of legal AI forward. We’re announcing 3 things:
1. RedlineBench with micro1: a benchmark for how frontier models handle complex, real world contract negotiations*
2. The Crosby Intelligence Research Fellowship: funding two Fellows pursuing frontier research with support from OpenAI - $25K & $12.5K Codex credits each
3. Hosting the most interesting conversations in applied AI at our SoHo office, featuring Parag Agrawal, Rahul Sengottuvelu, Peter Henderson, Neel Guha, Peyton Walters and more
So how close are frontier models to being able to do real commercial legal work?
We benchmarked frontier models on realistic contract negotiations (built externally with zero client data). Rather than individual edits, we focused on the full sequence of judgment calls a lawyer makes across a deal.
No model is close, and there are no standout winners yet.
The most interesting result from the benchmark is that every model is weakest on the opening move. You can read the full benchmark at the link in the comments.