transcribe

Your LLM Judge Is a Confident Liar: Building Better Verifiers — Browserbase

AI Engineer · 20m · transcribed 1h ago
More from AI Engineer Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Section Insights

# 0:00

Introduction to Browserbase and Research Overview

What is Browserbase and what research is being discussed?

Browserbase is an infrastructure company focused on deploying agents on the web, enabling web automation at scale. The research presented involves the development of a framework called 'stchen' and the evaluation of RL environments and language model judges as verifiers.

  • Browserbase specializes in cloud-based web automation.
  • The research focuses on improving agent capabilities through evaluation.
  • The framework 'stchen' is designed to ensure model improvements.
# 4:11

Identifying Gaps in Existing Verifiers

What are the weaknesses of existing verifiers?

Existing verifiers show a significant gap in agreement with human labels and often use less capable models. They lack proper rubrics and struggle with context management, leading to inaccuracies in verification.

  • Current verifiers have low agreement rates with human assessments.
  • They often use simpler models and lack essential evaluation rubrics.
  • Improved verification methods are necessary to close the accuracy gap.
# 8:23

Process Scores vs. Outcome Scores

How do process scores differ from outcome scores in evaluation?

Process scores assess the steps taken by the agent, allowing for partial credit for individual criteria, while outcome scores determine if the final goal was achieved. This distinction helps in identifying controllable versus uncontrollable failures.

  • Process scores provide a nuanced evaluation of agent performance.
  • Outcome scores focus solely on the final success of the task.
  • Separating controllable from uncontrollable failures enhances assessment accuracy.
# 12:35

Training with High-Quality Data

How does the quality of the verifier impact model training?

Using a high-quality verifier allows for filtering successful training trajectories, leading to the development of a higher quality model. This was validated through experiments showing that better verifiers yield better training data.

  • Higher quality verifiers lead to better training outcomes.
  • Filtering training data based on verifier performance is crucial.
  • The relationship between verifier quality and model success is significant.
# 16:47

Expanding the Universal Verifier's Application

Can the universal verifier be applied to tasks beyond web automation?

Yes, there are plans to adapt the universal verifier for desktop tasks, incorporating additional data sources like telemetry and logs to enhance its effectiveness in enterprise workflows.

  • The universal verifier is being adapted for desktop applications.
  • Incorporating more data types will improve verification accuracy.
  • Future benchmarks will include both web and enterprise task evaluations.

Transcript

0:12 Good afternoon everyone. thank you for coming. Hope you've had a great expo so far. We're in the last stretch, but I know it's been a wonderful, wonderful expo, at least for me. So hopefully you're learning a lot and we can teach you a little bit more. so I'm Miguel. I'm the tech lead for our agent platform at Browserbase. For those of you who don't know, browserbase is a infrastructure company for deploying agents on the web. it basically unlocks all of the web, not just what is easily accessible by managing browsers in the cloud and all the runtime that requires to do agentic web automation at scale. So let >> and I'm Corby. I'm a researcher at Microsoft and I've been collaborating with with browser base to help build really good verifiers.

1:02 >> So today we're here to talk about some research that we published about a month ago. and it is regarding evalance. It is regarding RL environments and LM judges as verifiers. So early on so I've been working on this open source framework called stchen for about a year and a half now. And from the beginning as any good agentic product needs to have we were building evals to make sure that the changes that we were introducing in the framework and that the model capabilities continue to improve and hill climb on those evals and at the beginning all those evals were deterministic environments. it was relatively simple. Model capabilities were some somewhat constrained and so we were able to build a lot of deterministic environments as model capabilities continue to improve as we continue to make improvements on the agent harness that just didn't scale. We found ourselves constantly patching or generating static sites coming up with longer trajectories with a lot of checkpoints to make sure that the evals were deterministically verifiable. And so there's many different project problems with automating the web at scale and verifying that what you intended to do actually succeeded. And it's that the web is very open-ended. There's not just one path to correctness. there is no ground truth. Sometimes the web changes and the product that you were checking for before is no longer there. So that breaks the whole determinism of your environment and there's blockages.

2:37 there's a lot of the environment errors that you face that don't let you get sort of a verifiable deterministic signal. So we recurred to what most people do which is trying to use an LLM judge to help improve and scale this operation. But we quickly found out that a lot of the leading benchmarks and some of you may know about OS World for computer use but there's some analogous benchmarks for web use called online minor web and web voyager.

3:10 Those come with LLM judges as verifiers. And what we noticed when we were using our sort of human expert annotators to verify the verifier is that in many instances they're very confidently wrong. And that leads to results that you can't really trust. So very quickly we realize that even if you were using this as an RL reward or as a signal to improve your hardness and auto research, you're not really training a better agent. You're just training a more competent liar.

3:47 And so in a real world use case at Microsoft, we trained this FARS 7B model which is a a web browser agent. And the same model on the same benchmark, we judged it according to the official web voyager judge, which is GPT40. and it said that it had 74% success rate. But there's a huge gap between the real truth which is when we used our new universal verifier which has high agreement with human labels that that number quickly becomes 30 like 38%. So there's a huge gap between what the existing verifiers say and what is actually the ground truth. And so we needed to find a way to to close this gap. The existing verifiers have lots of weaknesses. they're they use smaller and like much dumber models. as the LM as a judge they use 04 mini or GBT40 they don't even use rubrics so rubrics are the most critical thing that you need to have because you need to assign credit to where credit is due they don't look at the the relevant screenshots or they or they try to look at all the screenshots and they quickly get lost and overflow the context window of the LLM as a judge and sometimes they don't even look at the final answer or the action history of the model so our verifier checks all of these boxes and this is kind of how it works at a very high level you And you can look in the paper for more details, but given a task like book a cheapest flight from Seattle to Boston, the first thing we do is we generate a really good rubric and I'll give some examples of that. And a rubric has like maybe 10 different criteria of what success looks like. And then given that that criteria in the rubric, we look at the agents trajectory. All of the screenshots in this case are the the evidence of the ground truth. and we rank what are the most relevant screenshots in the trajectory for each criterion. And then we use that group of of top K evidence to determine whether that criterion was met or not and if there's any like differences or contradictions between what the agent said it did and what the state of the environment actually showed. And then given that we output a scored rubric which we we call like a process score because it gives partial credit in some cases. and we give an we output a an outcome boolean value to turn to basically say whether the agent accomplished the task according to what a reasonable user was would expect.

6:08 and so this is an example of what the rubric the rubric output is. It's basically a list of criteria as well as the outcome verifier is basically a true false flag with an explanation as to why a reasonable user would expect this trajectory to have succeeded or not. There were four guiding principles when we created this universal verifier. One is around rubric creation which is we wanted to grade only what was asked and not any extraneous criterion. we also did not want errors to cascade from one rubric criterion to another. We wanted them to be isolated. And I'll give an example of that. You also wanted to look at the ground truth screenshots.

6:50 Like this is a must. The agents will often overconfidently claim that they did something when in fact they did not do it. And so you have to look at the ground truth state. And then we also needed to separate what the agent could control versus what the agent could not control. The agent operates in a web browser. The web browser is an environment that it doesn't always have ability to control and I'll show some examples of that too. So this is an example of what a good and bad rubric looks like. The task here is to find a cheap hotel in Jakarta for these dates and then use the hotel's address to search for the closest coffee shop and then output the name and address of that coffee shop. Now a bad rubric which we did see happen would output a criteria saying like please tell me the total price for the stay at the hotel.

7:37 This is actually an extraneous criterion that was not asked for in the task but we saw a lot of rubrics naively generated do something like this and they would artificially deflate the scores because the agent didn't do something it wasn't asked to do. when it comes to not cascading errors, there are very subtle mistakes because a lot of tasks build on each other if they are like multi-step tasks. So in this example here, the task was to determine the net worth of the individual with the longest last name from among the members of NSync and Backstreet Boys. Okay. And so the agent in this case mistakenly thought that Timberlake had the longest last name which is only 10 letters whereas Kirkpatrick actually had the longest last name. So the model made one arithmetic mistake in one criteria of the rubric and that should not cascade to the next criteria of the rubric which is report the net worth of of the person that you chose. Right? So if you're able to report the net worth of Timberlake accurately then you shouldn't get penalized in this criteria. you should only get penalized in that one.

8:44 Hallucinations is the biggest one that we really wanted to focus on. So in this task, this is an example of a very subtle hallucination where you wanted to find some information about some image captioning model and the agent said that this model had plus 6.2% in some cider score, but in reality the abstract of the underlying paper said it was only 2.8% in cider score. And so this was a like a very subtle mistake that even humans will not catch unless an LLM will flag it. so this kind of gives an overview of of how process scores and outcome scores differ. in the case we mentioned one of the principles was we wanted to separate controllable versus uncontrollable failures. So, if the task is to buy, say, this plushy toy from Amazon, and the agent accurately searched for the plushy toy, but found that it was out of stock, it could not continue. And so, here we say that the agent gets full marks for doing its best effort to achieve the goal, but the outcome was still not met because it couldn't buy the thing that it wanted to buy because it was out of stock. so we basically enumerated a bunch of scenarios as to what could happen and how do we assign credit if the agent was able to find a similar alternative plushy toy in a different way. it would still get success there for both for both cases and it would get penalized if it made a controllable mistake. it would get penalized if it made hallucinations and so on and so forth. So we basically enumerated a a schema of like all the possible failure modes and how we would assign credit to that and the universal verifier adheres to those to those schema.

10:25 So >> basically in in order to evaluate the verifier and really tell whether or not we were able to hill climb the accuracy of that verifier, we started designing a lot of experiments with human experts and built a whole platform to collect data on what we call the ground labels that we've also released under cool verify bench for others to train upon. But the process that we followed was the following. At first, we would show the full trajectory of evidence to the human verifiers and the annotators would judge based on the evidence and the result of all the actions that the agent took.

11:05 After that, they would be prompted to a seeing the judgment of the universal verifier to say whether they agree or disagree with a human with a universal verifier. And we found this to be a very reliable signal to tell whether or not we were working in the right direction because I'll jump to this one straight up, but the verifier started correcting humans at one point. They it would identify things that the humans were missing. And so ultimately the result of the of the universal verifier on these golden labels was that we were able to reduce false positives from almost half to zero and coins cappa isometric to measure inner annotator inner annotator agreement. and what we can see here is that the h universal verifier agrees with humans as often as humans agree with one another. So we can see that 0.58 coins scalpa score versus the two annotators per task where we computed coins scappa as well to see how much they agreed with one another.

12:20 >> Another way another way to verify the another way to verify the quality of our verifier is to actually train on data filtered by it. So what we did was we did ran an experiment where we held constant the number of training examples. In this case it was 3k or 9k training trajectories, but we filtered those training trajectories based on whether our verifier said they were pass or not. And so when you're doing SFT, which is what we did here, you want to train only on SFT examples where they were true like successful trajectories. And so if you train on 3,000 trajectories filtered by the process score of our universal verifier, you will get a much higher quality model than if you were to train on a worse verifier like our own a baseline verifier. And this held at larger scales as well. And so this basically this experiment told us that if you have a higher quality verifier, it means you can filter higher quality data and training on that data leads to a higher quality model.

13:22 And so this was like the ultimate experimental proof in addition to the to the to the human agreements that this model this verifier is working and is and is reliable in a real production scenario. the last thing we did kind of as a side project is we wanted to determine whether AI in auto research can build the same verifier that we built. So basically what I did and what Miguel and I did is we sat down for three weeks to build the universal verifier and to tweak the prompts and to write the code for it.

13:56 And we basically ran about 30 different experiments over the course of three weeks to determine whether the verifier we were building agrees with human labels. Okay. And so that's what the Coen Kappa score here on the Y-axis is measuring. It measures whether an individual verifier system is agreeing with the human labels. And so the blue line here is the model Miguel and I, the system Miguel and I were were creating this auto research this universal verifier. And then we wanted to see if if we stripped everything that we did away, could AI in an auto research loop build the same verifier that we built with the same level of quality and fidelity? And that's what the red and green lines here are. The red line is if we took away all of our code and prompts, could AI replicate that? The conclusion that we got from this is that while we took about three weeks to build this verifier, an auto research loop could do it in about one day. It ran the same number of experiments in about one day. However, it only reached about 70% of the agreements that our verifier was able to reach. So there's still a gap there. And so the kind of the conclusion that I would draw here is that you can use auto research to build metrics and verifiers at least to help you speed up experimentation. but you probably still need some level of human intervention here. then again, these these results were from Opus 4.6. I haven't tried it with Fable yet. Maybe it will do better, but it's a pretty powerful baseline and it got pretty far along the way. So, not only can auto research help you build a model, but in this case, auto research can help you build a verifier as well, which we thought was a pretty cool thing. And and building verifiers is as important as building the models themselves is one of the takeaways I want to leave you with. So >> yeah, and just to double emphasize the green light is the auto research verifier primed with a lot of the findings that we had gotten from the three weeks of experimentation and in that optimization it was able to reach higher percentages than we were ever before. So combining that human intuition with auto research loops has proven to be the most powerful recipe for many research tasks. so the all of the work is open source. The paper is published under this as a preprint on archive. and the Microsoft repo contains the golden labels on kua verifier bench. It contains the code to run the experiments.

16:37 But beyond that this research also yielded an entire new benchmark. So we've talked briefly about online minor web and web voyager. I spend a lot of time working with labs helping make sure that we can provide them quantitative signal to improve their models. A lot of those open source benchmarks are very quickly getting saturated and it is very difficult to provide any sort of signal because it feels like the the test data is already in the distribution of the training data. And so this new benchmark has proven to have the biggest gap to complete success and it's my daily driver to make sure that I can tell whether model capabilities are strong, whether they lack and how to continue to close the gap. So >> did you have more slides?

17:28 >> No. >> Okay, >> we can take a little bit of I think questions, but we're good on time. >> Yes. Yeah. So the question is the question is whether we've used the the universal verifier for other kinds of tasks beyond just web web tasks. So one thing we're working on in Microsoft right now is to apply a version of this verifier for desktop tasks like things that you would do on your laptop like kind of enterprise workflows. And so we hope to publish another benchmark of not just web data but also like enterprise style desktop data. and we would have basically the similar the same kind of verifier for that too. It would look at a little bit more information because on desktops you have the terminal, you have more telemetry, more logs than just the browser. So the verifier would take those things into account as well. Yeah.

18:25 Any other questions? >> Yeah. >> So I personally never when you present I was thinking about like how do you think about that show?

19:00 Yeah. >> Yeah. So, the question is about how do you make sure you're not overfitting the verifier to the human labels because that is a problem. Miguel and I spent a lot of time working on this question. And so like one thing that we did was we kind of held out two sets of labels. One we we basically had a set of like 150 trajectories. on about 50 of those trajectories, I labeled them myself and I used them to hill climb in those 30 experiments. But for the other hundred trajectories, they were basically held out to us. They were done by humans that we had paid with like 2x overlap. So once our verifier was done being built, we gave it to those human annotators and each human labeled each trajectory, sorry, two humans labeled each trajectory and then that's what we published in KUA verifier bench. And so those are the numbers that we report here. so we're pretty confident that we didn't overfit because we held out at least twothirds of the labels. like we did not train like iterate on them.

20:08 >> Yeah. So that's very important is you don't want to train on like to be able to >> Yeah. Yeah. Yeah. The gold standard is is to hire humans and train them to do the the verification task themsel and then see whether your system agrees with it. Yeah. Any other questions? Anyone want to help us build more benchmarks?

20:43 >> Yeah. Oh yeah. Okay. Okay. Yeah. Yeah. Yeah. >> Awesome. Well, thanks a lot. hopefully you enjoyed the rest of the expo. It was a pleasure. >> Thank you so much. >> Thank you so much. >> Yeah. Yeah.

Summary

Miguel and Corby discuss their research on a new universal verifier for evaluating web automation agents, highlighting the limitations of existing verification methods. They emphasize the importance of creating robust rubrics and separating controllable from uncontrollable errors to improve the accuracy of agent evaluations.

- Browserbase provides infrastructure for web agent deployment, facilitating web automation at scale.
- The research focuses on improving verification methods for reinforcement learning environments using language model judges.
- Existing verifiers often produce unreliable results, leading to significant discrepancies between reported and actual success rates.
- The universal verifier incorporates a detailed rubric system to assess agent performance based on relevant criteria and evidence.
- Key principles include isolating errors, focusing on ground truth, and distinguishing between controllable and uncontrollable failures.
- Experiments show that the universal verifier significantly reduces false positives and aligns closely with human judgments.
- Auto research can assist in developing verifiers, but human insight remains crucial for achieving high-quality results.
- Future applications may extend the verifier's use to desktop tasks and enterprise workflows, enhancing its versatility.

Questions Answered

What is Browserbase and what research is being discussed?

Browserbase is an infrastructure company focused on deploying agents on the web, enabling web automation at scale. The research presented involves the development of a framework called 'stchen' and the evaluation of RL environments and language model judges as verifiers.

What are the weaknesses of existing verifiers?

Existing verifiers show a significant gap in agreement with human labels and often use less capable models. They lack proper rubrics and struggle with context management, leading to inaccuracies in verification.

How do process scores differ from outcome scores in evaluation?

Process scores assess the steps taken by the agent, allowing for partial credit for individual criteria, while outcome scores determine if the final goal was achieved. This distinction helps in identifying controllable versus uncontrollable failures.

How does the quality of the verifier impact model training?

Using a high-quality verifier allows for filtering successful training trajectories, leading to the development of a higher quality model. This was validated through experiments showing that better verifiers yield better training data.

Can the universal verifier be applied to tasks beyond web automation?

Yes, there are plans to adapt the universal verifier for desktop tasks, incorporating additional data sources like telemetry and logs to enhance its effectiveness in enterprise workflows.

© transcribe · For agents Built with care and craft by Gokul Rajaram