Mercor’s cover photo
Mercor

Mercor

Software Development

San Francisco, California 794,700 followers

Organizing human intelligence to power the AI economy.

About us

We find the best experts in every professional domain and put their knowledge to work training frontier models. Through APEX, we measure whether those models can actually perform economically valuable work. We're also bringing that expertise to enterprises: deploying custom AI agents, staffing teams with vetted domain experts, and helping organizations encode their own knowledge into AI systems.

Website
mercor.com
Industry
Software Development
Company size
201-500 employees
Headquarters
San Francisco, California
Type
Privately Held
Founded
2023

Locations

Employees at Mercor

Updates

  • View organization page for Mercor

    794,700 followers

    APEX-Agents score update for Muse Spark 1.1 When Meta first released Muse Spark 1.1, the model scored 37.1% Pass@1 on APEX-Agents, our agentic benchmark built for long-horizon professional services tasks across banking, law, and consulting. For roughly 10% of tasks, false positives in Meta's content moderation logic blocked trajectory completion, and as a result, those tasks failed and received 0-scores. We worked with Meta to help them update their content filters, which have since been reflected on the publicly available model checkpoint. The new result: 41.9% Pass@1, placing Muse Spark 1.1 at #2 overall, behind only Claude Fable 5 (43.3%). The update changes the APEX-Agents leaderboard substantially. Muse Spark 1.1 now ranks #1 in management consulting at 49.4%, significantly ahead of Fable 5 at 44.8%. Here's the full APEX-Agents domain breakdown: Management consulting: 43.9% → 49.4% Investment banking: 38.4% → 41.9% Corporate law: 28.9% → 34.5% Congratulations to the team at Meta AI on a great model release, landing firmly in the top tier of the APEX-Agents leaderboard.

    • No alternative text description for this image
  • Mercor reposted this

    There's no website that captures how a radiologist interprets a scan, how a banker sizes a deal, or how a lawyer builds an argument. That knowledge lives in people's heads. Mercor’s mission is to organize human intelligence to power the AI economy. Brendan Foody, Adarsh Hiremath, and Surya Midha built the network that finds the world's best experts, puts them to work training AI, and pays them for it, to the tune of $4M+ a day. From chess masters, radiologists, chefs, mathematicians … name a niche and they have it covered. Which is why nearly every top AI lab is a customer. We at Felicis are true believers in Mercor , the fastest startup to ever reach a $500 million revenue run rate. In March 2026, they reached $1 billion. And by June, Mercor had reached $2 billion in annualized revenue run-rate. Read the full story from Matt Quinn on how they went from college sophomores who started in a dorm room in 2023, to building core infrastructure for the AI economy today: 🔗 felicis.link/qPQCnhM cc: Sundeep Peechu James Detweiler

  • View organization page for Mercor

    794,700 followers

    We are opening new full-time roles across our SF, NYC, and London offices each week. Check out the 5 latest roles across EPD. 𝐑𝐋 𝐄𝐧𝐯𝐢𝐫𝐨𝐧𝐦𝐞𝐧𝐭𝐬 - 𝐅𝐮𝐥𝐥𝐬𝐭𝐚𝐜𝐤 𝐄𝐧𝐠𝐢𝐧𝐞𝐞𝐫 Build Studio, the internal platform that generates, reviews, and delivers data to the world's leading AI labs. 𝐀𝐬𝐬𝐞𝐬𝐬𝐦𝐞𝐧𝐭 - 𝐅𝐮𝐥𝐥𝐬𝐭𝐚𝐜𝐤 𝐄𝐧𝐠𝐢𝐧𝐞𝐞𝐫 & 𝐏𝐫𝐨𝐝𝐮𝐜𝐭 𝐄𝐧𝐠𝐢𝐧𝐞𝐞𝐫 Every expert who joins Mercor's network gets vetted, matched, and evaluated through our assessment platform. 𝐌𝐚𝐫𝐤𝐞𝐭𝐩𝐥𝐚𝐜𝐞 - 𝐒𝐨𝐟𝐭𝐰𝐚𝐫𝐞 𝐄𝐧𝐠𝐢𝐧𝐞𝐞𝐫 Search, matching, and allocation. The infrastructure that puts the right expert on the right project. 𝐂𝐨𝐫𝐞 𝐏𝐥𝐚𝐭𝐟𝐨𝐫𝐦 - 𝐒𝐨𝐟𝐭𝐰𝐚𝐫𝐞 𝐄𝐧𝐠𝐢𝐧𝐞𝐞𝐫 Own the shared architecture that every product team at Mercor builds on. Apply today at the link in the comments.

    • No alternative text description for this image
  • View organization page for Mercor

    794,700 followers

    Muse Spark 1.1 debuts at #6 overall on APEX-Agents, posting 37.1% Pass@1 and 52.0% mean criteria passed. It's a strong first showing, with one notable exception: for nearly 10% of tasks, the model introduced failures before completing the trajectory, and those tasks automatically score 0 on APEX-Agents. Much of that traces back to Meta's aggressive content filtering. Discount the affected trajectories and the pass rate climbs to 41.4%, moving Muse Spark 1.1 ahead of GPT-5.6 Sol. APEX-Agents is Mercor’s frontier benchmark for long-horizon cross-application tasks in professional services. This benchmark tests agents for real high-value work in investment banking, management consulting, and corporate law. Across the three APEX-Agents domains, here’s how Muse Spark 1.1 performed: Management consulting: 43.9% Investment banking: 38.4% Corporate law: 28.9% The model is also token-hungry. On APEX-Agents it averaged 1,550K tokens per task, behind only GPT-5.6 Sol (max) at 3,032K. Nearly 98% of that spend is input and context, with just 24.8K completion and 14.3K reasoning tokens. For comparison, GPT-5.6 Sol (xhigh) reaches essentially identical accuracy (37.7% vs 37.1%) on roughly half the tokens (757K vs 1,550K). Congratulations to the team at Meta on a strong debut.

    • No alternative text description for this image
  • View organization page for Mercor

    794,700 followers

    Grok 4.5 from SpaceXAI places #2 on the APEX-SWE leaderboard at 51.2% Pass@1 (±6.0), behind Fable 5 (65.5% ±6.2) on our benchmark for real-world software engineering work. It leads Integration (65.0% Pass@1) and places #2 in Observability (37.3% Pass@1), covering multi-step build tasks and diagnosis/debugging respectively. The Integration lead maps directly to the agentic workflows Grok 4.5 was built for: multi-step coding tasks run in collaboration with Cursor. Grok models have improved 30.2 pp in a year on this benchmark: Grok 4 (21.0% Pass@1) to Grok 4.5 (51.2% Pass@1). Congratulations to the xAI and Cursor teams.

    • No alternative text description for this image
  • View organization page for Mercor

    794,700 followers

    Earlier today, we published OpenAI GPT-5.6 Sol's APEX results. Now let's dive into what makes Sol distinctive: cost efficiency, how it achieves superior token economy, and what are the limits. Among frontier models, Sol (xHigh) delivers 49.8 Pass@1 points per million tokens on APEX-Agents. The next best is GPT-5.5 at 33.1. Fable 5 sits at 31.6, Grok 4.5 at 31.5, and Opus 4.8 at 30.2. That's roughly 1.5x the field. Sol reaches near-frontier accuracy on 757k tokens per trajectory while everyone else spends 0.9M to 1.4M. The mechanism is per-step leanness, not fewer steps. Sol averages about 28k tokens per tool call. Opus 4.8 averages 63k. Sol actually takes more steps than most of the frontier, but each one is less than half the weight. It works in many light moves rather than a few heavy ones. That is a different agent style, not just a smaller bill. So where does the efficiency run out? Quant modeling. On APEX-Agents, Sol's Pass@1 is 37.7%, compared with Fable 5 leading at 43.3%. The gap is not spread evenly across domains. Sol matches Opus 4.8 and GPT-5.5 in corporate law (60.2%) and management consulting (59.5%), but gives up 6 to 8 points to the frontier in investment banking (42.3%). Its headroom is concentrated in one place: heavy quantitative modeling work. There is another cost to running leaner steps: consistency. Given 4 attempts, Sol solved 48.5% of tasks (Pass@4). But it solved only 28.1% on all four attempts (Pass^4). That is a 20.4 point capability-to-reliability gap, wider than GPT-5.5's 15.0 points (45.8 vs 30.8). Sol can do more than its Pass@1 suggests. It just does not do it the same way twice. It is inexpensive and capable, but more variable across multiple runs. Finally, we wanted to see whether more reasoning spend would close any of these gaps, so we reran Sol on APEX-SWE at Max reasoning. Pass@1 rose from 39.7% to 41.2%. Integration held flat at 47.3%. Observability rose from 32.0% to 35.0%. These are real but slight gains, well inside the error bars. On APEX-Agents, Max effort actually scored lower than xHigh. Across both benchmarks the pattern holds: extra reasoning spend does not necessarily convert proportionally into higher scores.

    • APEX-Agents Model Efficiency Chart

GPT-5.6 Sol (xHigh) 49.8
GPT-5.5 (xHigh) 33.1
Fable 5 (Max) 31.6
Grok 4.5 (xHigh) 31.5
Opus 4.8 (Max) 30.2
  • Mercor reposted this

    I'm excited to announce that Mercor is acquiring Deeptune. While Mercor is the largest RL Environments company, Deeptune is the top vendor to Frontier Labs for the apps that make those environments work. The barrier to automating everything you do on your laptop with Claude or ChatGPT is covering the full distribution of apps, worlds, and tasks across the economy. Mercor's talent network has enabled us to lead the industry in scaling up the production of worlds and tasks. Deeptune is the best at the apps. Together, we make the best environments in the world. I couldn't be more thrilled to work with Tim Lupo and the entire Deeptune team.

  • Mercor reposted this

    Excited to share that Deeptune is joining Mercor. When we started Deeptune, we believed RL environments would become the defining bottleneck for AI. Over the past two years, we've seen this play out even faster than we expected. Mercor has built the world's leading expert network for AI, making it the obvious place to continue our mission. Together, we'll build the platform that powers the next generation of AI agents. Our entire team is joining Mercor in NYC, and we'll continue building with far greater resources towards the even bigger opportunity ahead. I'm deeply grateful to our customers, investors, and everyone who helped us get here. Most of all, thank you to the Deeptune team. None of this would have been possible without you. We're just getting started.

    • No alternative text description for this image
  • View organization page for Mercor

    794,700 followers

    Mercor is acquiring Deeptune. Training AI to do real professional work requires covering the full distribution of apps, worlds, and tasks across the economy. We have built the largest expert network and have been a leader in creating worlds and tasks. Deeptune built the apps, having created hundreds of them for the leading AI labs. Together, we can build environments across far more roles and industries than either company could alone. Every frontier lab is scaling environment production. Every enterprise will need to do the same. That is what we are building toward, and it is how we organize human intelligence to power the AI economy. Read our blog at the link in the comments.

    • No alternative text description for this image
  • View organization page for Mercor

    794,700 followers

    We evaluated Grok 4.5 on APEX-Agents, and it lands at #7 overall with a 34.2% Pass@1 and 50.9% mean score. That more than doubles Grok 4 on both metrics, a big improvement over the span of a year. APEX-Agents tests whether AI agents can execute the real day-to-day work of investment banking analysts, corporate lawyers, and management consultants. These are all long-horizon, cross-application tasks built and graded by industry professionals. Here’s how Grok 4.5 performed on APEX-Agents’ individual benchmarks: - Investment Banking: 37.3% Pass@1, 44.2% mean  - Corporate Law: 25.4% Pass@1, 52.2% mean  - Management Consulting: 39.9% Pass@1, 56.2% mean The new model ranked #5 in investment banking, while management consulting is its strongest domain. For corporate law, the gap between a 25.4% Pass@1 and a 52.2% mean suggests the model frequently completes most of a legal task but rarely delivers it completely correct, which is a big difference for professional work.

    • No alternative text description for this image

Similar pages

Browse jobs