Podcasts
Paul, Weiss Waking Up With AI
Technological Updates: Working Toward Safe and Structured AI Decision-Making
In this episode, Katherine Forrest and Scott Caravello discuss the White House Accord on Super Intelligence, a voluntary pact among major AI developers, before turning to two technical developments designed to help AI models reliably follow rules. They explore HardFlow, a technique proposed by MIT researchers intended to ensure flow-matching models used in areas like robotics take safe actions, and JEV, an AI model from TypeSafe AI designed to promote structured decision-making.
For the sources referenced in this episode, please see the links below:
HardFlow: Hard-Constrained Sampling for Flow-Matching Models via Trajectory Optimization
The White House: Inaugurating The Era of Super Intelligence
Episode Speakers
Episode Transcript
Katherine Forrest: Good morning, everyone, and welcome to Paul, Weiss Waking Up with AI. I am Katherine Forrest, and I want to say, Scott, before you introduce yourself, that I say it's morning, but I am going to guess based upon your background that it's not morning for you.
Scott Caravello: No, no, no, it is mid-afternoon. We've reversed roles from last week. I'm now in Europe; I'm in Sorrento with my dad. We're on a little bit of an Italian roots tour, so starting here and then going down to Sicily on Sunday, I think. Yeah. So it's great. But, you know, had to make time for podcasting. There's always podcasting.
Katherine Forrest: Fun, fun. You know what I love? I love the traveling podcast mic, right? It's such a great thing. You can, like, take this podcast mic, like, wherever.
Scott Caravello: Yeah, the only problem is that then I have to, like, just sort of scour wherever I am for enough, like, books or things to stack up in order to get a little podcasting mic to my, like, mouth height to speak into. So right now we have, like, a ton of travel books, like a Lonely Planet, sudden it'll guidebook, you know, we make it work.
Katherine Forrest: All right, great. Let's go to what we're going to talk about today in terms of AI. And as we all know, I think everybody knows, AI safety is all over the news. And we talked a couple of weeks ago about recursive self-improvement, which is AI that improves upon itself, and how that can sort of escalate and that being something people are watching out for. And we've talked about other sort of safety issues that are coming up. But today we're going to talk about AI safety from an entirely sort of different perspective. And we're going to start actually with a quick word on the voluntary pact between major AI developers and the White House that they had their lunch and then came out with. And then we'll talk about some technological developments that are focused on key questions of AI research and safety, which is, you know, looking for ways to get AI models to follow the rules. So that's where we're going.
Scott Caravello: OK. So I will kick it off with discussion of that safety pact, and then we can get to those two technical developments, which are really the heart of the discussion. But so, as you mentioned, Katherine, on September 29th, it was the White House brought in major AI companies and announced the White House Accord on Superintelligence. And so that included leaders of Anthropic, OpenAI, Google, Meta, NVIDIA and xAI, and they all signed the accord. And, you know, if you're wondering about that name, superintelligence, that is now the executive branch's official term for what we've been calling artificial intelligence. On the same day, actually, the president signed a separate executive order titled Inaugurating the Era of Superintelligence, which instructed federal agencies to say SI instead of AI.
Katherine Forrest: Right. You know, and that's actually interesting in and of itself. So we've got sort of two things happening here. First, let me just mention about this word SI versus AI, because I do think that this is not only a rebranding of artificial intelligence to something sort of like super, but I do think it's actually some acknowledgment that we're entering an era of superintelligence. But who knows? There wasn't really an explanation about why they're going with superintelligence now as a new rebranding. But anyway, on the accord, on this pact, it's voluntary, and it asks the companies, the developers of AI LLMs, to police themselves. And essentially each company agrees to put four layers of safety checks around its most powerful models, things like its own internal controls, an outside audit, a review by its board or a particular board. Again, voluntary, no penalties, but the companies do seem to be taking it very seriously.
Scott Caravello: Yeah. And the president had said that he only thinks the accord is, quote, morally binding. There is something to it. But so that's the policy piece. And I think we can move on to the first of the two technological developments that we wanted to discuss. The first is something called HardFlow, which comes out of MIT, and it's really, really interesting. And, you know, it's focused on what certain generative models produce and can do, but, you know, not really the LLMs that we typically use in our day-to-day life, but we'll get to that. But the idea is that this HardFlow technique can be used in safety-critical contexts where there just really isn't room for error. And so one more second before I turn it back over to you, Katherine, just to sort of flesh out that point about which generative models it's relevant to, because it's not those LLMs, it's for something that's actually called flow matching models. So that's where we get, you know, part of that term HardFlow, which are useful for, among other things, generating paths of movement. So, you know, most importantly, I think, for this context that we're thinking of, they're useful for robotics.
Katherine Forrest: Right. So let's pause on that for a second and talk about, like, flow matching models, because this is a new term for not only our audience, but it was something that I had to learn about by reading, you know, a couple of papers about it. And they're kind of a cousin of something we've talked about in prior episodes called the diffusion model, which is used frequently for sort of things like image generation and video. And both kinds of models work in the same basic way, both the flow matching model and the diffusion model. They don't produce a finished product in one shot. They start with what's essentially random noise, think of, like, static on an old TV screen, sort of like all that stuff sort of floating around, white noise. And then they refine it in multiple steps, so step by step, gradually shaping that noise into a more and more refined image, a finished image, or, to take the flow matching model, a finished path. So picture a sculptor starting with a rough block of stone and really slowly carving out the stone and ending up with an image or a statue. And you can get sort of the sense of this flow, you know, matching model.
Scott Caravello: And so I'm actually going to butt in really quickly, Katherine, to just say that everyone should keep that sculptor metaphor in mind because it will be important again. It's very, very useful.
Katherine Forrest: Right. And today, the usual way that we steer these models' processes is something that researchers call guidance. And at each step, you nudge the model more closely towards what you want, more or less. You know, you say things like, "I need more cat, less dog," and that kind of thing. But here's the crucial limitation with guidance: it's a suggestion, not a guarantee. And it makes the outcome that you want more likely, but it doesn't make it certain. Again, it's probabilistic. And so in some settings, more likely just isn't good enough.
Scott Caravello: Right, exactly. You know, because when we're talking about situations like physical AI where there's a robot involved, safety might be paramount. You don't want to only encourage a manufacturing robot or a humanoid robot to avoid an obstacle. You actually need it to do that. You need it to avoid the obstacle every single time, because the cost of failing to do so could be an injured person or a broken machine. And so that's what the hard constraints that this technique is putting on models, that's what it's designed to do. And so it's a rule that the model's output has to follow with no exceptions.
Katherine Forrest: Right. And the older way of enforcing that kind of rule had a real drawback, which is something worth understanding because it's really the reason for some of the academic work that's being done right now. The common approach is basically to yank the model, so to speak, into compliance at every single step of that generative process. But remember, the early steps are really a lot of noise, a lot of that white noise. And so you're forcing a fuzzy, half-formed draft to obey a final rule, and it's trying to, almost like going back to that statue sort of analogy, you're trying to fix the statue's nose while the rest of the block is still sort of an unformed block of stone.
Scott Caravello: There we go. There it is.
Katherine Forrest: Yeah. And so you end up overconstraining the whole thing, and the quality of the final output suffers. And so that's because in correcting the model at each step, you're actually moving it off the direction that you want it to take in order to generate your output.
Scott Caravello: Right. And so what HardFlow's really doing is getting to the right movement path or the right output in the most simplest and most reliable way that's also adhering to the safety requirements. And so that actually might mean those intermediate steps, the, the rough drafts along the way, letting that sort of not break the rules, but, you know, not trying to fix how it's complying every step of the way, because those drafts ultimately get thrown out. So only the finished product has to actually follow that rule.
Katherine Forrest: Right. So to close this out, the point is the rule only has to be met exactly at the very end, when the output is finished, and the messy middle of the generation process is left free for the model to explore, and in allowing the generation to roam while landing on a safe output. That provides an interesting aspect of the technique, because it's borrowing from an area of mathematics called optimal control, which is sort of the same kind of math used to plan a rocket's trajectory, and it uses it to steer the whole process. So the finished output obeys the rule exactly. In the researchers' tests on robotic arms and maze navigation and image editing, it met the constraints every time while producing better results than the older methods, like shorter, collision-free robot paths, which is a good thing because you don't want the robots to be running into you. And so that's tech development number one. So it's finding a new sort of architecture, new, if we can call it that, to find some safety ways of imposing some safety and guidance onto robotics. And so let's go on to the second tech development, which is actually a new kind of model altogether. And it comes from a San Francisco startup called TypeSafe AI, “type” and then “safe,” all one word, TypeSafe AI, founded by a former OpenAI researcher. And it's a model that they've named JEV, J-E-V.
Scott Caravello: Yeah. So the idea behind JEV is that the chatbot, right, that we're all so used to, a model that's generating text and other output, is great for talking to people, but it's not the right tool to use when you're trying to make a decision, whether it be AI making the decision or AI judging the decision made by another piece of software, which people do use more and more today because it's so smart and can do so many different things. And even though hallucination rates are down tremendously year over year for these LLM models, we do know that they can still hallucinate, and they can also try to give answers to users that they expect the users will like. So at least in TypeSafe's view, you know, it's not always the most accurate way to do this kind of judging and decision-making. And so if you're using AI to better guide an automated workflow within a business or another piece of software, that can cause issues.
Katherine Forrest: Right. So let's just pause on that, and sort of again, I just want to emphasize that this is moving away from our standard concept for the last couple of years of chatbots. That's really how we use these LLMs day-to-day, even in our enterprise environments, as well as whatever people are using when they just download it from the Google Play Store or the App Store. But now this JEV model really is starting to push things in a different direction. So let's talk a little bit about what JEV is and isn't. A normal language model, these LLMs that we talk about all the time, takes text in and puts text out, putting aside sort of the multimodal aspects. And then your software has to read that text and figure out what to do with it. Now, JEV, on the other hand, gives you a decision picked from a menu of options. You know, should this transaction be approved, yes or no? Which category does this document belong in? And the company calls it a frontier intelligence function call. OK, so that's, like, a big phrase, a frontier intelligence function call, which is a fancy way of saying input a whole bunch of messy information. Go ahead, put in a lot of messy information, but what you're going to get out is a clean, structured decision.
Scott Caravello: Yeah. And so why is that so interesting for this problem that they're trying to solve, right? Because first, like you mentioned, Katherine, the possible answers are fixed in advance. So you know what you're getting. The model can't just make up a new one. And it also can't hand back something to you that's garbled, you know, like, hedged, and you're not really sure what it's saying. It can only pick from the valid options. But of course, that doesn't mean that it's always going to get the answer right. A valid answer from the options you selected can still be a wrong answer, but at least it gives you something usable. But to that point about what's the right decision, every decision that it makes, that it generates, comes along with a confidence score. And so that's how sure the model is, and it's what it calls calibrated. And that's a really important word, because it means that if the model says that it's 90% sure across 1,000 decisions, it should be right about 90% of the time. Or that is at least what TypeSafe says. It is designed to tell you when it's actually unsure about the decision that it's making or recommending.
Katherine Forrest: Right. And that second feature is what really makes this JEV model usable as kind of a safety gate, which is what we're talking about alongside HardFlow. And you can think about it as an AI agent, or actually don't think about it, but think about an AI agent that we've talked about in past episodes. And, you know, an agent can take autonomous actions out in the world. And the big worry has been and is, you know, control. Can you actually prevent the AI agent from doing things that you don't want it to do? But a model like JEV can sit in front of the agent, the AI agent, as a kind of check, and before the agent does anything consequential, JEV can make a quick call. Is this allowed, yes or no? How confident is it that it's allowed, yes or no? How confident is it of sort of where it's going? And if the confidence is low, it gets to a human for review.
Scott Caravello: Which is pretty incredible. You think, at least theoretically, how it can sort of contribute to solving or mitigating a lot of problems that businesses have or think about when it comes to human review and making sure that they have eyes on the AI-enabled things and different other automated workflows that they want to have eyes on. But the sort of last thing I would add on JEV is that I don't have the stats right on hand, but that it's also supposed to be really fast and cheap. So it can actually do this over every single action and not just necessarily a sample, because it is so cost-efficient. And then finally, I think, Katherine, just to sort of add in that just this week, actually, at its big annual developer conference this week, but this is coming out next week, OpenAI announced something very similar, which it's calling the Decision API. Basically, it's a similar setup that you give a question and a fixed set of possible answers, and it gives back one of those answers along with a confidence score. And that's currently running on OpenAI's smallest and most affordable model, called Luna, and supposedly can generate back these decisions in about 150 milliseconds. You know, so it is roughly the length of a blink of an eye. So it seems to be that this idea may be catching on.
Katherine Forrest: Wow. It's actually, it's pretty incredible stuff. And so if we take a step back from both things that we've talked about today, there are caveats for these, these developments, both HardFlow and JEV. First of all, HardFlow, HardFlow, I have to get the words right, is a research paper. And while it's a strong result across several different areas, it's early. And so the testing has been in simulation. Nobody has really done a lot of reproduction of the test results yet. So we're going to see how that goes. And it's getting, you know, tested right now in various places. And so we'll see actually sort of where it ends up. JEV comes from a startup making its own performance claims, and we don't yet have independent data on how it's going to perform or what the results will be in the real world. So we're going to watch both of these things. But it's really, I think, useful as a takeaway to know and to understand, and your last point on Luna is part of this, that people are working really hard right now on safety. And Scott, that's all we've got time for today. I'm going to wish that you and your dad have great Italian food wherever you are.
Scott Caravello: Thank you.
Katherine Forrest: I'm Katherine Forrest.
Scott Caravello: And I'm Scott Caravello. Don't forget to like and subscribe.