transcribe

Loophole: Adversarial Agents To Stress Test Your Morality — Brendan Rappazzo, Morgan Stanley

AI Engineer · 16m · transcribed 3d ago
More from AI Engineer Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Section Insights

# 0:00

Introduction to the Project

What is the project about?

The project is an open-source game built on an adversarial agent framework that allows users to specify their morals and see how these can be codified into a legal system, while two agents test the system for contradictions.

  • The project is independent of the user's professional background.
  • It aims to explore moral reasoning through a game format.
  • Users can engage with their morals in a structured way.
# 3:20

Game Mechanics Overview

How does the game function?

Players specify their morals in natural language, which are then translated into a legal system. Two agents challenge this system by finding loopholes and overreach, while a judging agent assesses the outcomes.

  • The game operates in a loop of moral specification, legal drafting, and testing.
  • It utilizes LLMs for high-level moral reasoning.
  • Players can interactively refine their moral framework.
# 6:41

Example of Game Outcomes

What are the implications of the game outcomes?

The game can identify contradictions in a user's morals and suggest legal patches. For instance, it can highlight situations where something is legal but immoral, prompting user judgment.

  • The game reveals potential moral contradictions in legal systems.
  • Users can see the impact of their moral choices on legal outcomes.
  • It encourages critical thinking about morality and legality.
# 10:02

Practical Applications of the Game

What are the potential real-world applications of this project?

The project could help create moral frameworks for chatbots, facilitate decentralized contracts, and improve governmental efficiency by aligning laws with constituents' morals.

  • It can enhance chatbot interactions by embedding moral guidelines.
  • The project supports better contract negotiations by surfacing moral disagreements.
  • It envisions a more responsive government that aligns with public morals.
# 13:23

Future Directions and Government Efficiency

How can this project influence government processes?

The project could lead to more efficient governance by allowing constituents to measure how proposed bills align with their morals, potentially improving legislative outcomes.

  • It aims to quantify public agreement on legislation.
  • The project uses persona data to represent diverse moral perspectives.
  • It seeks to optimize legislative language for broader acceptance.

Transcript

0:12 I'll be talking about my project loophole. And I'm actually a machine learning researcher at Morgan Stanley, but this has nothing to do with Morgan Stanley. This is just a open-source project I've been building for fun. And to give sort of the high level flavor to start, it's really this game you can play that's built on top of this adversarial agent framework. So you specify your morals, one agent codifies that into a legal system and then these two adversarial agents try to find contradictions in your morals. And lately I've been building different extensions on top. but I wanted to, you know, start with sort of the origin story and and how I came up with this this idea. And so this really started, you know, a long time ago, I had sent my DNA into 23 and me for ancestry testing. and I kept hearing about, you know, more recently how DNA samples can be used, of course, to help solve crimes and all these forensics and cold cases. And I was thinking about how I had sort of opted out of of everything because, you know, I was scared of the kind of slippery slope and and how my DNA would be used.

1:26 but you know, there there are certain cases that I would be okay with. And it's sort of interesting. I was thinking like, you know, if someone could present to me case by case, you know, we'll use your DNA to solve, you know, help solve this cold case or this murder. I could sort of say yes or no. and I know where the the definition of like the and the nuance of my morals are. and but you know, of course, like enumerating all of this case by case is really sort of cognitively prohibitive. Like there's not a good way to do this currently.

2:00 And then I was thinking sort of more zoomed out that there's a lot of analogies to sort of the legal system as a whole. So, you know, one way to think of what a legal system is in our society is really just a way that we are trying to c codify our own moral beliefs. And I think sort of in a similar way like finding the true nuance of our morals and what the law should be as this really hard translation task. And I think often we kind of heir on the side of being too general because you know finding that nuance is really difficult and if you try to have a perfect translation it can we lead to these kind of weird corner cases or or weird failure modes and I even think kind of like peer-to-peer when we're relating to each other politically a lot of times the disagreements are more kind of fighting over core values which we don't really disagree on instead of exploring the really nuanced points of our our morals and I think you know following that kind of broader legal example I think in the you know the English common law system that's why we sort of lean on case law so heavily because we know that finding this this nuance and these nuance boundaries is difficult and so we kind of rely on smart judging to interpret and apply the law correctly and you know of course even with like the Supreme Court things can get elevated and we can decide whether a law is valid at all. and so that was sort of you know the idea is like can you take your morals and can you do this kind of synthetic case law generation. So you know this is overwhelming to do by hand but it seems like the new generation of LLMs are finally sort of smart enough to do this kind of highle moral reasoning and so that was sort of the the the starting point for this game. and I just want to take you through sort of the initial release of the game, the setup, how it works, and then also talk about some of the different branches I've been building on top of this open source project because I think it could go in some interesting directions. And so at a high level, how the game works is you in natural language and this all happens.

4:13 It's sort of like a terminal based game. You specify your morals and it could be your general morals or maybe about a specific subject and then there's one agent that takes those morals and drafts sort of a really rich legally codified legal system and then it just operates in this loop where one agent is instructed to try and find loopholes in your system. So something that is immoral but legal and another is prompted to find overreach. So things that are actually moral but illegal given your system. And a judging agent looks at the your morals the produced legal code and these sort of synthetic case law examples and first sees can it auto patch. So, like maybe the original draft of your your legal code sort of was an imperfect translation and there's not really a contradiction and it can just sort of autodo this update. Or maybe it's really kind of an underspecification of your morals or some kind of contradiction in your morals and in that case it raises it to you as the user to sort of be the judge and make a determination.

5:24 >> is it still on for you? It disappeared for me. Okay. so I know that, you know, if you I hope if you're curious about the game, you'll play it. It's all on GitHub. But I just wanted to show some examples. And this is a lot of text. So it's more about just showing the kind of shape of the the input and output. So this is sort of how you would provide your input. And going back to the DNA example, you might you know specify some number of of moral principles. And then the sort of codified legal system again just kind of looking at the shape has this really like legal ease, you know, preamble articles sections really trying to be you know write it in precise legal language. And then these the different kind of synthetic case laws get suggested. So in this case it's talking about this is a loophole it found where an insurance company trained a predicted machine learning model not on your DNA but on artifacts of the DNA. And so it's saying you know this is actually immoral but currently legal given your system. And in this case, it's it found that it could do sort of this auto patching and then you get this sort of like get style difference of of your original legal system and then the the difference it had to make to ensure you know this was consistent with your morals.

7:02 And then this is a an example of overreach and in this case it found that it couldn't do the auto patch. It's talking about, you know, someone submits their DNA for genetic research, but the researcher finds they have a rare but treatable genetic disorder. but currently your morals kind of say this shouldn't be allowed that they could disclose this disease to the the person submitting. And so this was raised to the user to me to kind of make a judgment. And then similarly, when you make the judgment, you get this this patch legal system. And so, you know, it's just sort of a a fun game and I posted on Twitter and shared it open source on GitHub. And for me at least, it was by far the most viral post I've had. And it sort of made me think like I think a lot of people just said it was sort of fun. You could stress test your morals, see if you have any interesting contradictions. But it also made me think, you know, is there maybe something more here? like could this be you know have more like practical or bigger scope implications and so I'll just talk about three different branches I'm kind of exploring the first and sort of leave leaving the legal area and really more practical is thinking about sort of an auto way to make constitutions for chat bots or really you know for agents in general where you know say you're a company and you want to have a agent or chatbot that's customerf facing and you want it to sort of adhere to a moral code but also have things it will and will not talk about.

8:35 I've kind of in one branch formulated it so you in a similar way write your morals. You write what the chatbot should and not talk about and then it tries to write this codified system prompt and then you have these kind of adversarial agents trying to get it to either talk about something it shouldn't or refuse to talk about something it should. And I see it as this sort of analogy or analogous method to GEA but really aimed at kind of building these codified system prompts.

9:07 The second use case that I'm I'm particularly interested in is thinking of it as a way to sort of do more ad hoc or decentralized contracts. So I think in in a simple case say like you can specify how you want your data or privacy to be handled online and you can go through this sort of adversarial game to get this codified legal system of how you want your your data handled and if you go to you know say Apple releases a new terms of service or something you can run the contradictions between your legal system and between Apple's terms of service and like surface any interesting contradictions or like synthetic cases where this would lead to a difference between how you know your morals, what you want and what the company is doing. And you know in the case that it's a big company, maybe you can't really change anything. It's not a negotiation, but you can at least be sort of have better information about the contract you're signing. But I also think in the case of you know thinking more decentralized like if you're trying to have contracts without you know some central authority kind of enforcing them and you're trying to maybe do contracts across different countries. thinking about like if you can specify your morals and how you want to like interface, you know, maybe it's just like contracted work, how you want your work to be paid for and and and the different morals surrounding that. And the other party can do the same. And then you both get this kind of stress tested codified contract. And then you can kind of find the the disagreements if there are any and surface them before you agree. And then you can kind of be more confident in the the contract as a whole.

10:58 And the last thing and maybe the kind of more aspirational angle is thinking about smarter government or more efficient government. I think there would be a lot of different privacy issues and logistical issues but sort of ignoring those for now and just thinking big picture. I think for voters or constituents, you know, this could be a really interesting way if you you defined your morals, you have this stress- tested legal code, sort of any new bill or politician that comes out, you could kind of run your contract against theirs and surface, you know, what are the the cases you would disagree or or interesting points that are kind of immoral to you or or a contradiction. I also think, you know, relating to one another, it's like a more I think we all have a lot of nuance in the way we feel about things and this is a way to kind of get to that nuance instead of arguing over just values which is, you know, often the values are not in contradiction. And then I think maybe a little more practically for legislators, you could imagine if you want to propose a bill and you can have like a a simulation of all the other legislators and a and a legislative body, you could sort of stress test it before submission. And so the the third branch I've been building on this project is I I tried to do this for the US Senate.

12:26 And so what I did is I first had Claude go through all current US senators and look at, you know, kind of all their voting history and anything else that was public and build their kind of moral system and then ran it through the loophole process to get a codified sort of legal code. And then on this system, you can, you know, take any current bill that's being proposed or even propose your own and submit it. And you can have Claude sort of simulate how each senator would vote.

12:57 And so here you can see like a breakdown of some senators, which way they're leaning and sort of their reasoning behind the vote. and I think you know it's sort of interesting just to think about like seeing what you know a proposed piece of legislation how people would vote but also this sort of becomes and I think on theme of the conference its own verifiable domain or loop and you could think about even kind of hill climbing the bill towards getting like a super majority or whatever you need it to pass. And so in this case, like this this Medicare bill I was testing, you know, it found that I think it originally started at like a 5050 vote and it found ways to hill climb the language of the bill such that it passed with 52 votes. And I think, you know, this is an example of it can find like the sort of the core tenants of the bill and it can try to find like run the the bill against each senator's contract and find is there any way I can change the language such that I don't violate sort of the core tenants or morals of the bill and kind of do those auto patching that way. And then it can also find you know kind of rank order the changes that would need to be in place to maximize votes and you as a user can kind of choose the trade-offs that way.

14:23 And then the last thing I I've been trying out more recently with this branch is actually looking at, you know, kind of even bigger picture, like can this lead to an even more efficient of government where you have every sort of constituent in a in a state or whatever the district is sort of have their legal code and then you could just submit any bill and actually measure sort of the agreement between like the the actual voters. And so for this, I took the Nvidia has this really great data set of USA personas. And so I took 500 personas per state and it's supposed to be sort of well representative of the state's population. Did the same process of having them given the persona, draft their morals, draft their sort of legal contract, and then take any bill you're interested in and kind of run it against each state. And you can also, you know, measure how much people like this bill or how much it's in agreement with their morals and then also do this hill climbing where you kind of optimize the bill for the people.

15:27 And so just to conclude, you know, at at minimum, I think it's a pretty fun game. I'm biased, but it's a lot of fun to just try out different you know, things you care about, put in your morals, see if there's any contradictions. you know, often it will raise some really interesting questions and then once you kind of provide that nuance, the game, you know, the the agents won't be able to find any more contradictions and you can kind of feel good that you have like a a consistent nuanced moral system. but I I am interested in, you know, exploring could this be are there kind of real applications here for some kind of like decentralized or better contracts and maybe even for legislators as a way to sort of stress test your bills and even think about how to write better laws that are, you know, better for the people in your district or more representative of what the people in your district want. and so this QR code is to the the Senate simulator. So I encourage you if you're interested to play and the other one is to my website which has the the full GitHub to loophole and please you know play with it fork it I'd love to have other contributors. Thank you.

Summary

The speaker discusses their open-source project, "loophole," which uses an adversarial agent framework to explore moral reasoning and legal systems. Users input their moral principles, which are then codified into a legal system that two agents test for contradictions, allowing users to refine their moral beliefs and understand the nuances of legal interpretations.

- The project originated from concerns about DNA privacy and the desire for a nuanced moral decision-making process.
- Users specify their morals, which are transformed into a legal system, while agents identify loopholes (immoral but legal) and overreach (moral but illegal).
- The game operates in a terminal interface, allowing for natural language input and output.
- The project has potential applications in creating ethical guidelines for chatbots, decentralized contracts, and smarter government legislation.
- The speaker has experimented with simulating U.S. Senate voting patterns based on senators' moral codes and proposed bills.
- The project aims to stress-test legislation and improve alignment with constituents' values.
- The speaker invites contributions and exploration of the project on GitHub, emphasizing its fun and educational aspects.

Questions Answered

What is the project about?

The project is an open-source game built on an adversarial agent framework that allows users to specify their morals and see how these can be codified into a legal system, while two agents test the system for contradictions.

How does the game function?

Players specify their morals in natural language, which are then translated into a legal system. Two agents challenge this system by finding loopholes and overreach, while a judging agent assesses the outcomes.

What are the implications of the game outcomes?

The game can identify contradictions in a user's morals and suggest legal patches. For instance, it can highlight situations where something is legal but immoral, prompting user judgment.

What are the potential real-world applications of this project?

The project could help create moral frameworks for chatbots, facilitate decentralized contracts, and improve governmental efficiency by aligning laws with constituents' morals.

How can this project influence government processes?

The project could lead to more efficient governance by allowing constituents to measure how proposed bills align with their morals, potentially improving legislative outcomes.

© transcribe · For agents Built with care and craft by Gokul Rajaram