RL Environments and Mercor's Data Market
Um, Record, I think you guys grew from a 1 to a 2 billion dollar revenue run rate in the last 4 months or so. Um, so this company's off to the races and I think you were just so front and center to how companies are thinking about uh post training their own models, uh building their own intelligence. So, thank you for joining us for this conversation. Um, format-wise what we're going to do is we've 15 minutes or so of content from Brendan. He's going to talk about uh RL environments in particular, which I think is a, you know, new frontier topic. It'll be fun to fun to explore. And then we're going to leave 15 minutes or so at the end for Q&A again. So, uh please keep please keep questions back pocket. I will turn it over to you, Brendan. >> Sweet. So, I'll be talking about RL environments. Starting out, I figured it's helpful to give a little bit of the background on the history of the data market and how that history ties into Record's origin story. Where things really started in 2020 in the era of crowdsourcing data for behavior cloning. So, this was mainly supervised fine-tuning data, inputs and outputs, and RLHF data where you would have a annotator select from a couple of model responses which they preferred. And we were able to make all this progress in fine-tuning GPT-3, making progress towards ChatGPT and GPT-4 in the crowdsourcing era of agentic data. But, what we saw changing in the market, especially as we uh headed into 2024, was this giant transition away from the low-skilled crowdsourcing era of behavior cloning data and moving towards the agentic era of data. Of how do we find the highest-skilled experts in the world that can work collaboratively in teams to build frontier evals and RL environments for the next generation of models๐1 models. All the software engineers, lawyers, doctors, bankers, et cetera that could measure the frontier of intelligence and help to use that to improve model capabilities. And so, Mercor grew up with our first big project being deep research. I guess the first prominent RL agent scaling up dramatically with all of the frontier labs to become the primary agentic data vendor to all of the leading labs and also all of the leading application layer companies ranging from Harvey, Cera, Cognition to Ramp. And what's been really exciting over the last 12 months especially is how RLVR within the agentic data paradigm has evolved to also include RL environments with these rich apps and worlds that teach agents how to use all of the tools on our laptops that we use every day๐1. So, I'll be talking about that and of course how this technology that started in the frontier labs is now getting disseminated to the application layer and all of the products that all of you are building as you work on your company. So, high level on what an RL environment is is that it includes three parts. The first part is the worlds. So, this includes all of the messages, slides, docs, sheets, etc. that correspond to everything you would have in a real project or company that you're working on. The second part is the apps which is high fidelity clones of popular applications, Salesforce, ServiceNow, Microsoft 365, etc. that agents can interact with via MCP, CLI, or Kua. And then the third part is the tasks where we have prompts and verifiers.๐1 Verifiers could be rubrics or unit tests that can be used either for eval or training.๐1 And the barrier for frontier labs to automate everything that you can do on your laptop using Claude is how do they cover the full distribution of all of the worlds, all of the apps, and all of the tasks in the economy. And so, there's been this enormous scale out in order to do that. Um, where humans have been really central to how we build these environments, obviously with models in the loop meaningfully. And so I put a graph here of the amount of expert hours that um, of throughput in from our talent network over the last 24 months. Um, and it's a a pretty crazy trajectory with respect to um, 2.5 million hours um, in Q2 alone with growth sort of accelerating on the amount of expert time uh, used to build out all of these environments. The reason being of course as I mentioned we need to scale out the environment distribution across every category in the economy. Many of you might know GDP val where there's 205 domains in the Bureau of Labor Statistics across all the different jobs, but then you have to think through how do we have all of the apps corresponding to all of those jobs, all the different scenarios, all the tasks. Is this enormous build out. Only humans can measure the frontier in most domains, not every domain. There are rare exceptions like math where you have a really clean simulation environment๐1 and so uh, the model's able to learn from whether it got the right answer, but in most domains like building a slide deck uh, the model has an incredibly hard time identifying reliably where it made its own mistake. It's as if you would be asking a human to grade their own homework. And so that's why it's really valuable to have a human create a rubric similar to how a professor would create a rubric to grade an essay or a TA would grade that slide deck. Similar to the way that a lot of us learn it's in large part from the feedback we got from those around us rather than uh, purely plugging things into a calculator or clean simulation. Um, and then building these verifiers is hard cuz anytime you're building the slide deck you need to understand the full problem space of what are the 10 different slide decks that, you know, could be a good path to go down? What are the dozens of mistakes you could possibly make? And how do you build a comprehensive verifier that captures this full solution area of what's possible? And so, what I'll walk through is a sample RL environment. Excuse me, to also break this down for all of you. Um and part of the reason that this is so cool, which I'll get to in a moment, is that we developed a lot of this technology in collaboration with the labs. These are, of course, ones that we have open-sourced and published to the world, but now that's all starting to get disseminated to the application layer companies that are building and owning their own intelligence๐1. As they realize that the three core pillars of their AI strategy are their compute, their algorithms or researchers, and the data sets they build. And data's often the most differentiating factor.๐1 And so, this is one that we published, um as a legal environment, where we have lawyers from top law firms like Latham and Watkins write out a scenario of a real project that they worked on in their big law job. And then they create a full outline for a data room that corresponds to all of the different uh messages, emails, files, size of files. I cut off the full data room cuz it's it's very extensive. Um and of course, there's a lot of model in the loop with how they effectively populate this. Similar to how a software engineer now should not be coding by hand entirely themselves, they should probably be orchestrating agents in how to do this very productively. Um and then we render that data room into the apps, the clones of Google Workspace you can see in this scenario, and have prompts uh to roll out model trajectories against this๐1. And so, in this one, it's evaluating the maximum total liability for Star Tanker uh Tankers International Limited compared to Cooper Jefferies Energy Corporation under the Oil and Petroleum Act, uh considering all of the context uh from this real scenario in the data room. And then as I mentioned, similar to how a professor would create a rubric to grade an essay, they have these key rubric criteria that correspond to what are the characteristics of a of a accurate model response. And making sure that these rubric criteria avoid reward hacking and effectively align with the uh when you roll out 100 trajectories, making sure all of those scores are accurate is incredibly technically challenging. And so there's an enormous amount of research, agenda quality control, training on the data, etc. that goes into how you solve that problem and then ultimately produce these high-quality verifiers and leaderboards um that give you an aggregate model score across how well all of the different models are um doing on a particular domain. And as we can see, one of the big changes over the last few months is that GLM 52 and Chimera K3 are on the leaderboard. And so that is a huge opportunity for all of you because that gives us the foundation to actually achieve frontier intelligence and all of the specific applications uh and verticals that you're focusing on um that's not too far away. And so to give a little bit of context on what that looks like, um I'll share an example of post-training on Apex Agents, which is the data set that I uh or the sample I just showed before, where we have 1,800 tasks in this example. This was post-training run of GLM 47, but we're redoing a lot of them for Chimera K3, so we'll have updated results for you all soon. Where you can see the jumps just on 1,800 tasks with about 500k in compute are pretty dramatic. Um corporate law going from 4.7% to 26.6%, um but notice that this is just Apex Agents data set we gave it, and it actually generalized incredibly well to GDP valve and uh Apex V1, which doesn't have these data rooms. Even just seeing nominal gains and some other benchmarks as well. And so, we're doing a lot of this work of working with customers like Harvey, who I know will present on stuff later to help build out the environments corresponding to their specific domain so that they can build frontier intelligence within that. And I think Andrew talked about how Cursor was a great first example of how an application layer company could build a industry-leading model that, you know, built an enormous amount of value for their customers. And I believe that over the next 12 months, there is going to be dozens of examples just like that where companies own their own intelligence and that is the key source of the modes that they're building. And actually, Josh and I talked about this the other day as well. Um a couple of examples of ways to curate high-quality data sets. The general three that we see most that I'm happy to talk about and send people links to is first by task is the most common. Where people would say, "I really like this data shape of environments in law and we will pay $2,000 per task to scale this up." And as an example, certain frontier labs might buy 50,000 tasks a month from us. And so, it tends to be pretty dramatic scale. And these tasks would generally be very complex. Some would even take humans up to a month to complete that given task. Sometimes it would take just a few hours. And these would be sort of custom per task pricing. Second is off-the-shelf data where we have we've invested hundreds of millions of dollars in building our own data sets that we sell to multiple customers. All these new neo labs are generally airing more towards off-the-shelf data because it doesn't make sense for 10 different labs to all be building their own uh data sets. Um there's a lot of value to building something once that can uh then be applied to everyone. And then the final which um we see a little bit of, but is uh less of our focus anymore, is just providing um the experts so that customers are able to uh or organize the experts on their own um in just an hourly model. Um so we we do a little bit of that when people like that was how Harvey got started with us hiring uh some lawyers, um but it generally moves towards more of these uh scaled offerings of data over time. So, that's a little bit of the background of how to build RL environments, what they are, and I'm really excited about all this technology that we have uh that has previously been limited to the frontier labs all making its way to all of you. And so, happy to answer any questions um about that. Sweet. Go ahead. >> Hey, uh um I'm Ali from Astro Guide. My question is it's kind of open-ended, but simple. How do you price data? Like how do you value data? >> So, there's a so many different ways. I mean, the most natural would be our customers care about model improvement, right? And so, our customers have a given goal of they want to you know, be at the frontier on a given leaderboard. And so, we're able to work backwards from how much is that worth to them and how much should we charge uh per task, how many tasks do we think would get them to that goal. And so, when we think about uh a company like Nvidia, they're probably willing to pay, you know, a billion dollars to have a frontier open-source model. And so, there's a lot of complexity of like how do we you know, price all the different ingredients that go into um making that happen. The other way that we price when it lens we look at it through is also our cost structure to make them where of course when we have a task that takes 10 hours of human time and we're paying the human $150 an hour there might be a $1500 cost basis and so then it becomes a question of what margin do we want to run on top of that based on how differentiated and frontier that specific task is. But it's it's super wide range. Like we have tasks that range from $50 to $10,000. So. >> How do you think about data quality? You mentioned like you know utilize human expert to label data and how do you how do you compare the human label data and you know the frontier lab you know frontier alliance the judgment data? How do you compare them to from your opinion? >> So first question was how do we think about quality? Second one was sort of how do we compare the like judgment of preference labels to the auto graders? >> Yeah. >> So the core spot which is generally when people say data quality they're referring to two things. First is realism and secondly is accuracy of verifiers.๐1 On realism it's just like they want to automate everything in the economy that corresponds to corporate law in this case right? And so it's like how do we make sure that this actually reflects the real distribution of what we would see in a real lawyer's environment. And that's one of the reasons that experts create outlines and help to guide all the processes of the data curation. Realism of the environment the apps the tasks everything is incredibly important and and also granularly understanding the taxonomy that drives that realism across the entire distribution that you're looking for. The second part of it relates to the other way people think about quality, which is the accuracy of the verifiers. Cuz the way that you would train one of these models is you might roll out 100 trajectories of Kimi K3 and then use this rubric to score all of those trajectories๐1. And as you can imagine, there's like so many different paths that a model can go down. And so you want to make sure that the way this rubric is doing the scoring is the same as if we were to just have human stack rank those 100 trajectories. And so what we do for that is a process called trajectory analysis, where we roll out 10 trajectories of the model that we're focused on improving, um, and then score all of those, and have some combination of agentic quality control systems and some human review go through to make sure that, um, all of the scores align with, uh, the goals๐1. Um, and sometimes you can also use, uh, human feedback evals or preference labels to as a eval for your auto grader, um, is the other way related to that that you're able to solve for it. Go ahead. >> Um, how much do you think like synthetic data generation plays into all of this, especially like creating these large data rooms? >> So the fascinating thing is I think that there's been a lot of misinterpretation of what people mean when they say synthetic data because, like, RLVR is a bet on synthetic data. It's basically let's roll out a bunch of synthetic model trajectories rather than having the humans write the SFT, and then let's score all of them, and let the models learn from all of these like, uh, synthetic model trajectories.๐1 So I think that's the first way that synthetics get used. The second way is that models play giant role in the way that we populate environments and create tasks in the same way that, um, a lawyer that is writing a legal memo should definitely be using Claude or ChatGPT to do that. The experts that are building out these data rooms should definitely be using Claude, ChatGPT, or whatever model to help them do that. And there's a lot of ways that the model can make them more efficient. But the reason that humans are still an essential component of the process that's incredibly differentiated is that you need humans almost definitionally to measure what is beyond the frontier of the model capabilities. Like the models you you can't just tell the model like come up with the legal environment and then like tell me which of your legal memos are like good and bad. It's super noisy and there's not like clear signals associated with that. You need something that has capabilities beyond the frontier of that model to do so reliably. >> Thank you for doing this. Um my question is RL environments seem like they're all the rage now and maybe have been for about a year. I had I hadn't really been hearing about them prior to that and it was all human expert labeling. And so kind of why why why is it all about RL environments now? Is that the most relevant thing for application companies to be thinking about? And is there any something after RL environments? >> Um so I'll I'll start with why it's become the rage and maybe some of the differences also between the deep research paradigm and the like envi- the sort of environments paradigm as we saw it in 2025. And then I'll talk about looking forward what we see evolving in the data landscape. Specifically, I think that the reason the deep research environments were the first was because like deep research had tool use with search. So search was the tool in the environment that the model would work in. But the experts would not necessarily be populating apps. So it was sort of a lighter version of an RL environment where they would just like create rubrics corresponding to this. And again, like I only talk about this stuff cuz it's a couple of years old at this point, so it's no longer uh super confidential. And then for the like trend of apps in 2025, I think that really became giant because people realized that the primary bottleneck to making the models useful was how they started to use both all the context in the code base and all of the tools on our laptops๐1, right? And so if we want this in the user distribution of usage, then we need to get it in the data distribution that the models are learning from.๐1 Uh and so there's going to continue be to be this giant scale up of diversity across all three of these categories on a going forward basis, but there's going to be some changes um to maybe name two of those changes that we're thinking about the most. The first one is ultra long horizon. Like right now agents mostly aren't trained to do things that are over 10 hours, and we need to start building tasks for things that might take a human 100 hours or even 1,000 hours to do๐1. And so that's going to be a giant shift. And then the other large shift that we're seeing is introducing virtual co-workers๐1, which corresponds to that. Um like one my favorite questions to ask people when they're thinking about their data distribution is what percentage of tasks that they do in their job require interacting with other people. Uh and most people would say like 60% or 70%. Some people say a lot more, some people say a little bit less. Um but then if you map that on to what percentage of evals measure how well the models can interact with other people, it's like 1%, maybe Tau bench has a little bit of this. Um and so there's this giant realism gap associated with how you actually measure how well agents engage in social interaction um throughout uh all of the different people and other agents that they need to work with uh in their jobs. Good. >> You talked about rubric generation, which is like is it like a bespoke access of verify that you put task? And from what I understood, that's like bottleneck by experts. Have you found any success of like being able to scale that up with your models giving you like some heuristic or even like post any models for >> We have found that you can make it a lot more efficient if you have a like AI copilot that's able to work with the expert in creating the task and the verifier. Um so, the expert can talk to the trajectory and understand exactly what's happening and where it's going wrong.๐1 The challenge is just that if you're trying to improve FABLE, FABLE cannot reliably write out the like rubric criteria for where it's making mistakes. It might get like half of them right and half of them wrong, and that amount of noise is unworkable from a training standpoint. Um and so, that's the reason that the the process that requires humans the most is the task creation. Uh like a lot of the environments, we can get like use a lot of synthetic. Uh it's helpful to have humans write the outlines they're familiar with the environment and and uh and they're grounded in reality, a realistic distribution, but with the task that's those tend to really require humans. Um with with rare exceptions, um in code or if you're sort of distilling from like if you have a model that's worse than Kimikaze 3, then it can definitely learn from tasks Kimikaze 3 is creating. So, there are there are some exceptions if you're doing it that way. >> Hey, this is uh Nikhil from uh Cribl. So, when we think about RL environments for certain provable domains like cyber defense or incident response, where there's the model is or the agent is trying to find a flaw in an existing system, Do you use uh humans for just authoring that environment or setting it up or do you also use that for grading? Is there Is there a way to scale that up? >> I I actually think cyber is one where you don't necessarily always need humans for the verifiers cuz you can have an attacker and a defender agent๐1. Um and uh I think you're right in saying that for cyber you can have humans more so or or architect what is a realistic like environment um and sort of set up the environment um cuz you do need a lot of diversity. Uh but then it's less human intensive with respect to uh building verifiers. >> Thanks. >> How can you tell when you're limited by the base model? >> What do you mean by that? >> Well, ostensibly you're using the same data set for all these models here and they somewhat land around the same final performance on this list here. But maybe if you try to smaller model, which is maybe a good place to start, it would land much lower. Um what is the cause there? Is it just parameter count or >> So the parameter count will definitely play a role in how effectively the model does insofar as how trainable it is. I think that um the main thing to look at is generally the gap between the like pass at 16 and the pass at one๐1. If you have a model where you roll out 16 trajectories and it gets all of them totally wrong, the it's sort of like hopeless that the model is going to learn from that um for the most part. Maybe you roll out another 100 trajectories and it gets one of them right. Um versus if you have um the ideal case is that you have um pass at one it fails, but then pass at 16 when you roll out 16 trajectories it gets it right once or twice, and then the model is able to learn very effectively from that๐1. So, that's generally the heuristic we would use for how strong the base model needs to be to effectively learn from a given data set. Cool. Maybe Oh, final question really quick. >> What advice do you have for companies as they partner with you on data that they should rely on themselves as part of their RL post-training versus relying What's the complementary >> Well, I think this is why a giant portion of our business is custom data, where it's like we have teams that are siloed and fully exclusive to critical customers to make sure that we build out the best data sets in the world that they own. And that allows them to maintain their competitive advantage๐1 associated with this, while also benefiting from all of the infrastructure that we've built. And I think there are some companies that try to like build out all of the talent network and infrastructure in-house, but I think if you look at what the frontier labs do and the best models do, it's a pretty good indication that there's so many economies of scale from working with a partner that has all of these economies of scale, the platform, the talent network, etc. So, anyways, thanks for all having me. >> [applause]