Transcript
Speaker 1: Bloomberg Audio Studios, Podcasts, Radio News. Speaker 2: Hello and welcome to another episode of the Odd Lots podcast. Speaker 2: I'm Jill Wisenthal, and I'm Tracy all the way, Tracy. Speaker 2: One thing about AI is that lots of lines that Speaker 2: go up. Speaker 3: Yes, famously, there is perhaps one line that has captured Speaker 3: the attention more than others when it comes to lines Speaker 3: going up. Speaker 2: Yes, but we're recording as April seven. Did you see Speaker 2: the anthropic revenue chart, by the way, Oh. Speaker 4: It's just extreme. Speaker 2: Okay, it's just the number of lines going up. I Speaker 2: mean there are some really. Speaker 3: Let me caveat that. Up until recently, there was one Speaker 3: chart of a line going up exponentially that became I Speaker 3: think it's fair to say the most viral chart in AI. Speaker 4: Right, Yes, I would absolutely agree with that. Speaker 2: So one of the many lines that go up, or Speaker 2: there are various lines that sort of capture This is Speaker 2: essentially just measures of AI progress of what they could do, Speaker 2: what the models are capable of, and so forth. And Speaker 2: you know, there's all different benchmarks out there and hobbyist Speaker 2: benchmark creators, et cetera, all kinds of benchmarks out there. Speaker 2: Organization called Meter based out in San Francisco, and they Speaker 2: measure how well AI models are doing at various sort Speaker 2: of engineering tasks, et cetera. And they have these charts Speaker 2: showing how long, you know, certain tasks, how long would Speaker 2: take a human to do them, and then whether AI Speaker 2: could do them. And yes, the lines just almost vertical. Speaker 2: I think there was someone one of the ones that Speaker 2: came out maybe very early this year or late last year, Speaker 2: showing the latest Claude model. Speaker 4: It, yes, like this is crazy. Speaker 3: When I look at these charts, they're called time horizon charts. Speaker 3: When I look at them, I like, intuitively I kind Speaker 3: of understand what they're saying, and you can kind of Speaker 3: see the leap in progress between some of the previous Speaker 3: models and Claude, right the latest claud model. And that's Speaker 3: what got everyone excited, was you had this big exponential Speaker 3: shift up in the capability of that particular AI model. Speaker 3: But then when I start like diving into what it Speaker 3: actually says on Meter's website about what these charts represent, Speaker 3: I start getting really confused. I know everyone wants to Speaker 3: get excited about AI and charts going up in general, Speaker 3: but I think there's a lot of nuance here and Speaker 3: we should probably talk about it, because the other thing Speaker 3: going on with meter right now is they've become sort Speaker 3: of the industry standard benchmark, and so a lot of Speaker 3: investment decisions are being based on these charts. And if Speaker 3: you oversimplify them as just like okay, lines going up Speaker 3: and then suddenly it goes up even more, obviously, people Speaker 3: are going to start to get like maybe a little Speaker 3: over excited. Speaker 2: Can I see one other thing too that I'm very Speaker 2: curious about, Like, I'm really glad that there are people Speaker 2: designing various benchmarks for measuring a progress. Seems like an Speaker 2: important thing to get a handle on. But like if Speaker 2: I were, like, say, like talented or smart enough to Speaker 2: be like doing these things, I would go work for Speaker 2: one of the labs and make ten million dollars a Speaker 2: year or something like that. And so I'm actually curious Speaker 2: to a lot of the nonprofits, et cetera. It's like, Speaker 2: do you really want to be like working at the Speaker 2: cutting edge of AI in a nonprofit? I mean, I Speaker 2: guess Open Eyes owned by a nonprofit weirdly enough, but Speaker 2: you know what I'm saying, Like I would want the money. Speaker 3: We should talk about it with our guests who are Speaker 3: currently sitting right here. Speaker 4: That's exactly right. Speaker 2: I'm very excited to say we have the two perfect Speaker 2: guests to talk about the best of viral and maybe Speaker 2: important chart in AI. Right now, we're going to be Speaker 2: speaking with Joel Becker. He is a member of the Speaker 2: technical staff at METER, and we're also going to be Speaker 2: speaking with Chris Painter, the president of METER. So, uh, Speaker 2: Joel and Chris, thank you so much for coming on oud, Speaker 2: lots for having us. Yeah, really excited to chat with Speaker 2: both of you. Chris, and you're the president. I'll start Speaker 2: with you, like, what is METER? How long has it Speaker 2: been around? What is this organization? What's its goal? Just Speaker 2: give us the sort of six sixty second synopsis of Meter. Speaker 5: Yeah, totally. Speaker 6: I can try and you know, sometimes I give a Speaker 6: long version. I can try and do a short version here. Speaker 6: So Meter is a research nonprofit based in the Bay Area, Speaker 6: like you said, dedicated to advancing the science of measuring Speaker 6: whether and when AI systems might pose catastrophic risks to Speaker 6: humanity as a whole, focused specifically on threats that come Speaker 6: from AI autonomy or AI systems themselves. So when you Speaker 6: talk about there's kind of this whole field and AI Speaker 6: of dangerous capability evaluations. People seeing ken this AI system Speaker 6: assist with a chemical or biological weapon attack, can it advance? Speaker 6: Kind of like bad actors' ability to execute cyber attacks Speaker 6: on a really large scale. METER is sort of specialized Speaker 6: in specifically assessing how autonomous are AI systems, what is Speaker 6: the scale and like length and difficulty of tasks that Speaker 6: they're able to do by themselves, partially because we think Speaker 6: it sets the stakes for conversations about AI misalignment. So Speaker 6: we sort of see ourselve being on the hook for Speaker 6: at any given point in time, giving humanity the bits Speaker 6: of evidence that are most informative for establishing the stakes Speaker 6: of are we reliant on AI systems as a society Speaker 6: in a way that could make it really bad if Speaker 6: they are misaligned. Speaker 3: I'm going to let Joe ask the question about why Speaker 3: you're both working in a nonprofit instead of one of Speaker 3: the labs later, But one question I do have is Speaker 3: when I think of METER, you guys always come up Speaker 3: in the context of these time horizon charts. And I Speaker 3: don't mean this as an insult or anything, but I Speaker 3: hardly ever hear anyone talk about the actual safety aspect Speaker 3: of your mission. Why do you think that is. Speaker 6: Yeah, So I think there's some distinction between our motive Speaker 6: for assessing time horizons and the kind of how it Speaker 6: gets used then by the rest of the world or Speaker 6: kind of like what the origin of the rest of Speaker 6: the world's interest in it for meter. I think the Speaker 6: reason that we work on things like the time horizon Speaker 6: charts is because if we're trying to establish the stakes Speaker 6: for talking about could AI systems go rogue or one Speaker 6: day could they like try to take over and subvert Speaker 6: human control? Three years ago, if you went back to Speaker 6: around when it meters started about fourish years ago, and Speaker 6: if you it was started by Beth Barnes Paul Cristiano Speaker 6: and this was kind of the initial motive. Is if Speaker 6: you went back then and you said, why don't I Speaker 6: think that AI systems are going to go rogue and Speaker 6: like take over or overthrow humanity today the kind of Speaker 6: most intuitive you know, you can come up with a Speaker 6: lot of abstract reasons debates about the goals AI systems Speaker 6: might or might not eventually have, but the kind of Speaker 6: most damning in the moment reason is the AI system Speaker 6: just can't do much right. It doesn't make sense to Speaker 6: talk about a question answer system that like can't even Speaker 6: reliably answer programming questions saying like is it going to Speaker 6: hack my systems or like backdoor me in some way. Speaker 6: It just doesn't make any sense to talk about. Speaker 3: It's going to write you a poem that you asked. Speaker 6: Right, or won't even at the time they couldn't do Speaker 6: anything for themselves. And so if you're like, kind of Speaker 6: being able to subvert human control depends on agency, And Speaker 6: so we wanted to come up with a measure that Speaker 6: kind of tracks agency over time to kind of say, Speaker 6: when would this argument no longer? When are AI systems Speaker 6: now able to kind of do long, complex enough actions Speaker 6: by themselves that the argument kind of the goalposts almost Speaker 6: move somewhere else to like, well, we would catch the Speaker 6: AIS or the AIS don't want to subvert human control. Speaker 6: And so I agree that there is a distinction between Speaker 6: like how I think partially the exercise of trying to Speaker 6: come up with these measures throws off things that are Speaker 6: very like grounded and intuitive measures of AI progress that Speaker 6: might be more intuitive than just benchmarks. Right, So if Speaker 6: you a lot of people are in the game of Speaker 6: making just benchmarks, where you say, like here's my harm Speaker 6: bench or something the AI gets seventy percent. That's much Speaker 6: less of a kind of grounded or long lasting metric. Speaker 6: Like it's hard to say what that means or how Speaker 6: that generalizes, but the idea with time horizon is like, Speaker 6: maybe it's more intuitive, and I think that helps both Speaker 6: for safety and for like business understanding. Speaker 4: So let's talk about what this charge. Speaker 2: I go the main chart here at meter dot org, Speaker 2: right on the front page, it's this time horizon chart, Speaker 2: and it shows Claude Opus four point six as a Speaker 2: February twenty twenty six able to complete a task length Speaker 2: in eleven hours and fifty nine minutes with a fifty Speaker 2: percent success rate. I have to admit, the first time Speaker 2: I saw this chart or version of this chart, what Speaker 2: I assume, and I suspect others assume, is that it Speaker 2: was able to go off and work on a task Speaker 2: for eleven hours and fifty nine minutes then come back Speaker 2: with an answers. But apparently it's not that. What do Speaker 2: you walk us through what's really being measured here? By Speaker 2: the way, the previous high was all was GPT five Speaker 2: point three codex. That was five hours of fifty minutes. Speaker 2: So I guess part of the reason this charge is Speaker 2: blew people in mind because literally that's basically a double Speaker 2: But why don't you talk to us about what's really Speaker 2: being measured here? Speaker 7: Yeah, so fundamentally, you know, in simpler terms, we are Speaker 7: plotting the difficulty of tasks the AIS are able to Speaker 7: complete over time. And you know, the particular way that Speaker 7: we measure the difficulty of tasks is in how long Speaker 7: it takes humans to complete those same tasks that we're Speaker 7: asking the AIS to do. So in this case, you know, Speaker 7: talking about for a PUS four point six, something like, Speaker 7: tasks that take humans twelve hours to do, we predict Speaker 7: that it will succeed at those tasks around fifty percent Speaker 7: of the time. And yeah, you know, it turns out Speaker 7: that when you plot using this particular difficulty measure how Speaker 7: performance AIS are relative to how long it takes humans Speaker 7: to complete these tasks, we see an exponential increase in Speaker 7: capabilities for AIS. And what that ends up meaning is Speaker 7: that you keep on having these doublings of capabilities every Speaker 7: let's say four months. Speaker 5: It seems on recent trends. Speaker 7: Where you know, the next model is not merely going Speaker 7: to have necessarily, you know, an hour longer time horizon, Speaker 7: but perhaps be having some multiple of the. Speaker 5: Time horizon of the previous model that's come out. Speaker 2: So then explain how the number of their twelve hours Speaker 2: is established. So there is some engineering task and you say, okay, Speaker 2: this is a test that would require twelve hours, but Speaker 2: humans of all different types of talent capabilities, how do Speaker 2: you establish that? Okay, this was a twelve hour TESK, Speaker 2: this was a six hour tel whatever. Speaker 7: Yeah, So the simple answer is, literally, we get humans Speaker 7: to sit down and complete the tasks that we give Speaker 7: to AIS and as close to identical conditions as possible. Speaker 7: So first we come up with the tasks, and that's Speaker 7: you know, that's whole god a killer fish. We can Speaker 7: talk about exactly how we do that. And then, using Speaker 7: essentially the same tools that we're about to give the AIS, Speaker 7: we take talented humans, you know, not people who have Speaker 7: seen this particular type of task before, but people who Speaker 7: have relevant expertise. So if it's a software engineering task, Speaker 7: you know, they have software engineering expertise. Machine learning task, Speaker 7: they have machine learning expertise, and then we time them, Speaker 7: we see how long it takes for them to complete Speaker 7: those tasks cescfully and then roughly we call the difficulty Speaker 7: of the task as measured in human time to complete, Speaker 7: as the average time it took these humans to complete Speaker 7: the task. Then we'll run the AIS on this same Speaker 7: set of tasks. Typically today for the very easiest tasks Speaker 7: that they're more or less always going to succeed, there's Speaker 7: some mid range of tasks where you know, perhaps they Speaker 7: succeed fifty percent of the time, or perhaps for some Speaker 7: tasks in that range they succeed zero percent of the Speaker 7: time and for others one hundred percent of the time. Speaker 7: And so they're getting fifty percent on average, let's say, Speaker 7: and then for the much harder tasks, perhaps they're getting Speaker 7: closer to zero percent. And then the point of which Speaker 7: we predict, you know, in the middle of all these Speaker 7: zero percents and one hundred percents by task, the point Speaker 7: at which we predict that they'd have a fifty percent Speaker 7: chance of succeeding. That is, either a fifty percent chance Speaker 7: of succeeding on some task or fifty percent of the Speaker 7: tasks or of that's difficulty that we think they would Speaker 7: succeed on. That's what we're going to call the time Speaker 7: horizon of these models. Speaker 6: I think one thing also that could be good to explain. Speaker 6: Here is the task distribution. I mean, this is not Speaker 6: this is not all activities that humans do. We are Speaker 6: specifically here interested in or the like. There's some question Speaker 6: in what tasks are you know, like Joel mentioned, we're Speaker 6: having people come into our office do the task to Speaker 6: get a sense of how long it takes. We're not Speaker 6: having them come in and like, you know, paint paintings Speaker 6: or write novels or you know. We're focused here specifically Speaker 6: on things that are in the distribution of work that Speaker 6: a engineer at a like. We like to think of Speaker 6: it as like a frontier AI lab, the tasks that Speaker 6: they might be doing. So this is things like software engineering. Speaker 6: It's fine tuning at AI models, it is like software Speaker 6: machine learning, that kind of task. Speaker 3: Wait, can I just ask why did you decide to Speaker 3: focus on engineering? Because you could have winded out to Speaker 3: you know, if we're talking about AI being capable of, Speaker 3: you know, taking over the world, there are all sorts Speaker 3: of substantive tasks that would fall under that category. So Speaker 3: why just do engineering? Speaker 6: Yeah, I think that for one thing, maybe other people Speaker 6: in the team or maybe Jolis thoughts about this, but Speaker 6: I think my particular motive and being interested in the Speaker 6: time horizon on software tasks is that first of all, Speaker 6: it's the thing that the industry is very like already Speaker 6: even before we started working on this, is very focused on. Speaker 6: So it's one of the capabilities that you should expect Speaker 6: to come along for the ride earliest. It's the thing Speaker 6: that like a lot of optimization pressure is being exerted on. Speaker 6: And then I think that it is kind of the Speaker 6: like thing that you would expect as an early warning Speaker 6: kind of sign of this AR and D automation. So Speaker 6: to some extent, METER thinks of itself as trying to Speaker 6: build you know, science that are advanced science that can Speaker 6: say when are we getting to the point that aisism Speaker 6: could improve themselves or speed up the pace of AI development. Speaker 6: When will AI research kind of feed on itself? And Speaker 6: the kind of core capability for that might be software Speaker 6: engineering and machine le learning research ability. There are other Speaker 6: skills that could be relevant to taking over the world. Speaker 6: I think other people have done time rising some like Speaker 6: cybersecurity sense. Speaker 3: But I suppose it is true like the basilisk isn't Speaker 3: going to paint its way into like power or something Speaker 3: like that. Okay, it might. Speaker 2: Deceive you it might be very convincing or cunning in Speaker 2: some way, and handover the cues. Speaker 5: I always say for your mental models. Speaker 7: You know, we don't have perfect evidence of this whatsoever, Speaker 7: but my rough sense, sort of colloquially or you know, Speaker 7: my prior before evidence comes in, is that if we Speaker 7: did study tasks on these very different distributions, you know, Speaker 7: not machine learning, not software engineering. I'm not sure about Speaker 7: painting exactly, but you know, perhaps or other kinds of Speaker 7: task distributions that we could enumerate that basically we would Speaker 7: see this similarly shaped exponential progress over time where every Speaker 7: I'm not sure exactly, but let's say, you know, for month, Speaker 7: six months, something like that, the level of capabilities as Speaker 7: measured in time horizon would be doubling at something like Speaker 7: that pace, maybe from a much lower level. So you know, Speaker 7: one example that we do have better evidence of is Speaker 7: that the ais today are much less performance at you know, Speaker 7: anything that requires vision capabilities, seeing what's on a screen, Speaker 7: clicking around at a computer, but they're getting you know, Speaker 7: tremendously better that sort of thing over time. Speaker 5: I just do mention quickly. Speaker 6: We did actually do a very kind of brief investigation Speaker 6: of this another task distributions that's on our website somewhere Speaker 6: like cross domain time horizons. I think we looked at Speaker 6: data from the Tesla's shared on self driving and forgetting Speaker 6: the other there's like os world. Maybe some of these Speaker 6: are like somewhat similar, still kind of in the distribution Speaker 6: of software tasks, but trying to get further afield into Speaker 6: things like vision. Speaker 3: How big is the sample size on the humans who Speaker 3: are actually doing work? And also is it getting harder Speaker 3: getting like human engineers into the room to compete with Speaker 3: like Claude Opus four point six versus say, if I Speaker 3: was a mediocre engineer, and I'm not, I'm a non Speaker 3: existent engineer, but if I was a mediocre one, I Speaker 3: would like maybe I would feel good about going up Speaker 3: against like GPT three or something, and maybe I would Speaker 3: feel a lot worse about myself going up against like Claude. Speaker 7: Yeah, you know, on these tasks, I'm in a pretty Speaker 7: similar position myself to you. So we have approximately three, Speaker 7: although it varies quite a lot across tasks. Human baselines Speaker 7: per tasks, so you know, typically we're ever going over Speaker 7: something like three. I think the final numbers, it's my Speaker 7: impression that they're not going to be so sensitive to Speaker 7: the particular baselines that we use. Speaker 6: Aren't the longer tests week more weekly baselined. Speaker 7: Yeah, So indeed, I think it will get a lot Speaker 7: harder to baseline these tasks as the length of task Speaker 7: AIS are able to successfully complete gets longer and longer. Speaker 7: You know, you might think at some points the length Speaker 7: of task that they can complete is longer than the Speaker 7: time in four months time, they're going to be able Speaker 7: to complete tasks of more than four months, and then Speaker 7: it's you know, kind of becomes paths close to impossible Speaker 7: to get these four months long baselines. Of course we're Speaker 7: not at that point yet, but you know, definitely has Speaker 7: become more difficult to get these baselines as time has Speaker 7: gone on. Speaker 5: At the moments, not impossible, but very challenging. Joe. Speaker 3: These are the future jobs for displaced engineers, right. It's Speaker 3: competing against the codes for benchmark first benchmark evaluation, we Speaker 3: found the jobs. Speaker 2: So we mentioned at the beginning the most viral CHARLINEI Speaker 2: is this chart that you have on the front of Speaker 2: your website. Your website defaults to this and it shows Speaker 2: you know, this doubling. So if we actually go back Speaker 2: to November let's say November twenty twenty five, Gemini three Speaker 2: pro three hours and forty four minutes, claud Op was Speaker 2: four point six twelve hours. Those are the fifty percent Speaker 2: success benchmark. If we go to the eighty percent benchmark, Speaker 2: which the website doesn't default to improve the price of Speaker 2: improvement looks a little less impressive to me. So okay, Speaker 2: now it's like it does not have the same gap Speaker 2: pretty clearly. Now eighty percent is still not one hundred percent. Speaker 2: And I know that this is your meter's goal is about, Speaker 2: like you know, human safety and all this stuff. But Speaker 2: when we think about people look at this and they Speaker 2: use it as a stand in for how performance are Speaker 2: these models? Even eighty percent, you know, certainly for like Speaker 2: any business application. I understand you're not like serving business Speaker 2: here per se, but probably businesses care about this. Even Speaker 2: eighty percent may not be very good enough. And it Speaker 2: does not look as crazy when you look at the Speaker 2: eighty percent chart as it does at the fifty percent chart. Speaker 2: Why the focus on the fifty percent chart? And given like, Speaker 2: why not look at the chart that just does not Speaker 2: look as impressive. Speaker 5: Yeah, maybe two central things to say. Speaker 7: One to my eyes, the eight percent shot it's basically Speaker 7: does look as impressive. Well, the doubling time is about Speaker 7: the safe cope. Speaker 5: On it's the same. Speaker 6: It's the same, it's say an afseet of it's the Speaker 6: same pace of progress. Speaker 7: You know, it's something like five times smaller than the Speaker 7: fifty percent than the fifty percent number. Speaker 5: But you know that only takes you too doublings. Speaker 7: And if each doubling takes around four months, that means Speaker 7: that in eight months time you're going to have the Speaker 7: same eighty percent success rate roughly as you do fifty Speaker 7: percent success ray today. Speaker 5: That's one thing to say. Speaker 7: Maybe a second thing to say is, you know, remember Speaker 7: at the beginning I said, essentially what we're doing is Speaker 7: plotting the difficulty of tasks that these ais can complete Speaker 7: over time, just with this particular measure that ends up Speaker 7: showing this clean exponential trend. And we've picked a particular Speaker 7: number as our difficulty number, and you know that is Speaker 7: this fifty percent reliability threshold. We could have picked a Speaker 7: different one. I think there are reasons for picking the Speaker 7: fifty percent one. In particular, it's the one that statistically Speaker 7: we're better able to measure. For some technical reasons, it's Speaker 7: the one that shows up in previously. It's show that Speaker 7: there are some couple of a couple of other reasons Speaker 7: why we can go for fifty per cents rather than Speaker 7: eighty percent. Maybe a final thing to say is that Speaker 7: this fifty percent number is sort of equivocating between these tasks. Speaker 7: It's able to complete fifty percent of the time and Speaker 7: fifty percent of the tasks it's able to complete one Speaker 7: hundred percent of the time, and fifty percent it's able Speaker 7: to complete zero percent of the time. And actually, I Speaker 7: think the situation is it's somewhere in between, but it's Speaker 7: a little bit closer to the latter, where there are Speaker 7: some tasks that it's completing with near perfect reliability and Speaker 7: some tasks in that range that it's completing with very Speaker 7: low reliability. And for downstream economic applications or for applications Speaker 7: inside of these maju AI companies or something, you know, Speaker 7: you might think that that's more favorable in some sense, Speaker 7: that there are some of these tasks where we're getting Speaker 7: one hundred percent reliability, even even for very challenging tasks. Speaker 6: I think two other things make maybe it could be Speaker 6: useful to just explain when you said that they are Speaker 6: technical reasons why it's easiest to measure it fifty percent. One, Like, Speaker 6: it is just the case that it is fifty percent Speaker 6: is the point at which it is like least sensitive Speaker 6: to Like the distribution is kind of thickest, right, I mean, Speaker 6: correct me if this is wrong, But my I mean Speaker 6: there are like to resolve something like ninety five percent, Speaker 6: you would need way more samples because then you need Speaker 6: to have some that are like you need way more Speaker 6: simples to be able to resolve that level of precision. Speaker 7: I think there are some caveats to that picture. But Speaker 7: let's say even more extreme. You know, let's say that Speaker 7: we cared about you know, ninety nine percents. In that case, Speaker 7: if we had one percent label noise quotes unquote, yeah Speaker 7: it is you know, if sometimes we were accidentally grading Speaker 7: some of the failing tasks, passing some of the passing Speaker 7: tasks as failing, then we just never be able to Speaker 7: estimate that reliably right, And at fifty percents, this comes Speaker 7: a little bit closer to washing out. Speaker 6: And I think one other intuitive thing here, or one intuition, Speaker 6: is that if you give me a task and you Speaker 6: give me the model, it is the point at which Speaker 6: I think that the model, all you tell me is Speaker 6: the time or the length of task that it takes Speaker 6: a human to do the task. The fifty percent time Speaker 6: horizon is the point at which I think it is Speaker 6: more likely that the model will be able to do Speaker 6: the task than that it can't. And I just find Speaker 6: that intuitive. Speaker 3: Yeah, how much interest do you get on these charts Speaker 3: from potential investors specifically? And the reason I ask is Speaker 3: because I was just messing around and like googling some Speaker 3: stuff and when the OPUS chart, the latest opis chart Speaker 3: came up, someone posted it on Breddit and I think Speaker 3: like the second comment on it was someone going, how Speaker 3: do I invest in open Ai? And like and like Speaker 3: people were they were trying to club together to like Speaker 3: invest in these companies. So clearly there are people out Speaker 3: there who are using these charts as investment tools. Speaker 6: I would say, you know, we don't get an enormous Speaker 6: amount of inbound from investment firms. I mean, sometimes, you know, Speaker 6: vcs or whatever we're based in the Bay area will Speaker 6: reach out to us. I think that there's some kind Speaker 6: of principle of our goal is to inform the public Speaker 6: and give them the best evidence that we can about Speaker 6: when we might get to this point of kind of Speaker 6: you know, AI being you know, fully autonomous or able Speaker 6: to improve itself. And there's some principle at play here Speaker 6: of like I kind of want to enable people to Speaker 6: do whatever they will do with that information, and I Speaker 6: think that we don't engage a ton in kind of Speaker 6: the like business side or investment implication of the work. Speaker 6: One kind of thought experiment I sometimes say to myself Speaker 6: is if I do believe that at some point we're Speaker 6: going to get this AI that's improving itself, and where Speaker 6: like AI research is automated, and you have all these Speaker 6: fears about a singularity, would I rather that like all Speaker 6: of Wall Street like falsely didn't think that was coming Speaker 6: when I believed it was coming, Or would I want Speaker 6: them all to know that it was coming, given that Speaker 6: I believe it's coming, and I think all of human Speaker 6: Maybe this is more a personal view, but I think Speaker 6: if this is possible that we will automate AI research, Speaker 6: I think all of humanity being aware of it, aware Speaker 6: of where we're heading, is sort of a precondition for Speaker 6: us all being able to figure out what to do Speaker 6: about it. And so I don't kind of want like Speaker 6: certain people or one side or one team to kind Speaker 6: of like selectively be in the dark because they might Speaker 6: invest on the basis of this or something like that. Speaker 6: But we don't, you know, it's not where we put Speaker 6: our time. We're focused on informing the public. The public Speaker 6: includes some investors. Speaker 3: So on that note, like what is the actual level Speaker 3: at which we're all presumably supposed to panic or at which, like, Speaker 3: if you're a policy maker, you would start to get Speaker 3: worried about AI being able to automate and improve on Speaker 3: itself in a way that eventually becomes detrimental to humanity. Speaker 7: I don't know exactly what the level is on this time, Speaker 7: horise and measure. I think, you know, one thing to Speaker 7: say is we have made real progress on the science Speaker 7: of measuring these AI systems and how capable they are. Speaker 7: But I think there's a long way to go, and Speaker 7: in an important sense, I think we're behind on this task. Speaker 7: We're measuring some underlying technical trend and at some point Speaker 7: I do think that implies greater risks. Speaker 5: So astonishing things happening. Speaker 7: Although Chris can speak more to other arguments that we Speaker 7: might back out to for why even if AIS are Speaker 7: very capable, we still might not see castrophic dangers emerge Speaker 7: in the short term. Speaker 5: Yeah, I'm sure, you know. Speaker 2: I think part of the reason why the AGI cheddar Speaker 2: has really picked up, particularly in the wake of like Speaker 2: everyone using Chlord code, is it's very easy to emerge Speaker 2: in it. So like you're sitting there, it's like, yeah, Speaker 2: do this, do this. It's like I don't even need Speaker 2: to be here, right, I think you sort of get Speaker 2: a very intuitive feel for like how the human could Speaker 2: come out of the loop. What helps today, because I'm Speaker 2: sure there's been tried. Like if you go to like Speaker 2: ch juput and you say, here's a you hear you Speaker 2: have cloud code access, go build something, and the AI Speaker 2: is what actually happens today when AI is working with A. Speaker 7: Yeah, my sense is that at some point, you know, Speaker 7: further away points than would have been true some time ago, Speaker 7: the AIS will more or less full on their faces Speaker 7: that you know, there are some things they're not so Speaker 7: capable of today, like collaborative hallucinations. Speaker 4: Will they're just like, you know, just like devolved terribles. Speaker 5: Yeah, I think all sorts of ways can go. Speaker 7: You know, at some point they're going to need to Speaker 7: rely on external resources, and today that they're not as Speaker 7: capable at managing these external resources effectively. I think they're Speaker 7: less capable at sort of ideation and sort of self Speaker 7: awareness about where they are in the problem today than Speaker 7: they are at these kind of raw software engineering skills. Speaker 7: You know, you know, as you mentioned, the ways in Speaker 7: which AIS are autonomous today or close to autonomous today, Speaker 7: is the human has the idea and then you know, Speaker 7: submits that idea to cloud code or a code or Speaker 7: one of these other agenticie tools, and then they handle Speaker 7: the software engineering components and possibly there's still still some Speaker 7: intervention after that. I do imagine that the sort of Speaker 7: circle of autonomy or something gets larger over time. I Speaker 7: do think there's no fundamental barrier. It seems to me Speaker 7: today as having those ideas and so we moved to Speaker 7: a great level of abstraction. But if we were purely Speaker 7: relying today on these fully autonomous capabilities, you know, could Speaker 7: you manage research departments, any any apartments of your choice Speaker 7: inside of a major AI company. Speaker 5: No, my guess is probably not. Speaker 3: Actually on this note, this reminds me something I wanted Speaker 3: to ask. So when you look at the domain specific Speaker 3: time horizon charts, so the ones that show like I Speaker 3: think you call them task suites or something like that. Speaker 3: Like I guess productivity by a specific job, and you Speaker 3: see these different lines, So sometimes you see like almost Speaker 3: horizontal lines and sometimes you see squiggly or steeper lines. Speaker 3: What is actually happening there? Like, how are we supposed Speaker 3: to interpret that? Like is this a measurement problem? Or Speaker 3: is it say something very fundamental about like what AI Speaker 3: can and can't do under current conditions. Speaker 6: The thing that I think would be good for Jill Speaker 6: to explain is that I think that there is a Speaker 6: distinction here between will AI like the time horizon charge Speaker 6: doesn't by itself, I think, tell you will productivity in Speaker 6: one specific kind of job increase because of access to AI? Speaker 7: Yeah, maybe one thing to say on that chart showing Speaker 7: the time horizon on these different task distributions relative to Speaker 7: my guesses ahead of time, You know, I think those Speaker 7: time horizons are remarkably similar. I think the doubling times Speaker 7: the pace of progress in AI seems more similar than Speaker 7: I put of guessed to the original trend that we Speaker 7: that we published, although you know, imperfectly so on this Speaker 7: difficulty translating what we might call raw AI capabilities, in Speaker 7: some sense you know, capabilities on benchmarks or something to Speaker 7: real world productivity. I think there are a number of Speaker 7: differences in a number of ways, in particular in which Speaker 7: the benchmark results are overestimating what we might see in Speaker 7: the wild, you know, not not hugely overestimating. I think Speaker 7: we do see that people are getting real utility out Speaker 7: of these modern agentic a IOLs, but overestimating to some extent. Speaker 5: One is that the. Speaker 7: Scoring implicitly is different in real problems, I'm scoring based Speaker 7: on something a bit more holistic than these algorithmic scoring procedures, Speaker 7: these automatic scoring procedures that we're using at Meter and Speaker 7: many other people that are using in the in the Speaker 7: benchmark world. There's some notion of code quality if you're Speaker 7: if you're working in software engineering, but for other tasks Speaker 7: there's there's. Speaker 3: Beautiful code, elegant code. Speaker 5: People talk that yeah, yeah, yeah, for other tasks that's Speaker 5: going to be coding. Speaker 1: This is. Speaker 7: One more thing is that the tasks that come up Speaker 7: in the wild are more likely to be messy in Speaker 7: some sense. They involve working with other people, They involve Speaker 7: working in much larger code bases or sort of more Speaker 7: open ended problems, maybe with something even adversarial going on Speaker 7: in the in the software engineering context, that might be Speaker 7: that someone's trying to make a change to the part Speaker 7: of the code base that you're currently working on and Speaker 7: you need to and you need to work around that, Speaker 7: and we do tend to see that the AIS are Speaker 7: less capable working on these more messy problems. Speaker 5: I don't want to overstate that. Speaker 7: You know, it's not an enormous effect, but you know, Speaker 7: that's one thing that gets in the way of these Speaker 7: productivity increases, you know. And then I do think there's Speaker 7: something to the reliability question right where you know, if Speaker 7: it was true that for a certain type of task Speaker 7: you only had you know, eighty percent reliability, then every Speaker 7: time you're going to need to go back and verify Speaker 7: the work of these AIS, and not only verify the Speaker 7: work of dcais, but without the context of how they Speaker 7: implemented the solution relative to if you went about the Speaker 7: task yourself, you'd already have that in your head, and Speaker 7: so this verification step quote unquote would take less time. Speaker 7: You know, I don't expect these frictions to be sort Speaker 7: of so fundamental in some sense, or I imagine they Speaker 7: go up levels of abstraction I think not only as Speaker 7: the underlying technical progress real, but I think that the Speaker 7: productivity improvements that are also going to show up increasingly. Speaker 5: But yeah, there are these frictions. Speaker 2: Tracy alluded to this question when she asked about VCS Speaker 2: and investor interest. So people see these charts and regardless Speaker 2: of what Meter's point. Speaker 4: Is, like, this is incredible. I got to invest in this. Speaker 2: But this brings me to this broader thing that I Speaker 2: find very strange about AI, which is this kind of odd, Speaker 2: sort of Baptist and bootlegger relationship between the AI labs Speaker 2: people who are building this stuff and the sort of Speaker 2: alignment safety people, and they sort of go back and forth, Speaker 2: and like you have the heads of the lab saying yes, Speaker 2: this might destroy the world and take all your jobs, Speaker 2: and the safety people in the alignment people. Speaker 4: Says, yes, this might destroy the world. Speaker 2: And like, I'm very strange industry right, Like the only Speaker 2: thing that I can think of a cigarettes, where like Speaker 2: they warn you that smokey is bad, except they had Speaker 2: to do that because they lost a lawsuit. Speaker 4: I don't think they were particularly inclined to do that. Speaker 2: I can't think of any other industry where the most Speaker 2: enthusiastic people about it are also warning and dooming about Speaker 2: how bad the thing they're building could be. So I'm Speaker 2: sort of curious, like you know, first of all, like Speaker 2: and I talked about this in the intro, like who Speaker 2: is the type of person that's like working it meter? Speaker 2: It is like skilled enough to do like advanced evaluations, Speaker 2: and like where's the funding coming from? But like talk Speaker 2: to us about like who's behind meter and why they're there. Speaker 6: Yeah, totally, So. I think one thing to say on Speaker 6: the history of kind of people caring about AI safety Speaker 6: in the day area is that this concern goes back Speaker 6: like quite a ways, I could say for over a decade. Speaker 6: There are many people who got into the field because Speaker 6: they saw this trend of deep learning, Like what if Speaker 6: deep learning works and it kind of goes all the Speaker 6: way to artificial general intelligence and then superintelligence, and if Speaker 6: that works, then it could affect everything. I think possibly Speaker 6: when people worry about this, there's a future that they Speaker 6: have in mind with super intelligence that's even more capable Speaker 6: than what people who think of themselves as like AGI Speaker 6: pill today think of. They're imagining AI systems that can Speaker 6: run you know, the entire economy and I think people Speaker 6: who kind of a while ago or many years ago Speaker 6: saw that vision and were sort of alarmed about the Speaker 6: stakes of it. Many people had this intuition that the Speaker 6: thing to do is go and work in the industry Speaker 6: because if you're like helping build it, you know what's Speaker 6: the best way to shape the future, It's to build it. Speaker 6: And I think that there's obviously you could have questions Speaker 6: about how sincere that is for many of the people Speaker 6: who are in the industry, or if there's kind of Speaker 6: a mix of different motivations and like you know, different Speaker 6: wolves inside of them where maybe they partially are motivated Speaker 6: by that, but also they're like there's kind of this Speaker 6: like Oppenheimer, like, it feels good to feel like you're Speaker 6: in the position of making something that's dangerous made. Speaker 2: Someone wants described Open aiyet of me, this is years ago. Speaker 2: Friend said it was like Open AI I was sort Speaker 2: of like the Manhattan Project, except the goal was to Speaker 2: not build the bomb at the very end, if that Speaker 2: makes any sense. So to your Oppenheimer point, it's like Speaker 2: very strange. Speaker 6: And I think one thing to emphasize is, you know, well, Speaker 6: it could be that there's a mix of motivations now Speaker 6: there are definitely many people, I think in the Bay Speaker 6: Area who sincerely believe that the technology is headed to Speaker 6: someplace that will be very difficult for a huge where Speaker 6: it will be very difficult for humanity to stay kind Speaker 6: of in the driver's seat or like stay in control Speaker 6: and kind of a meaningful sense. Speaker 2: It does seem is though, like people talk about all Speaker 2: the big AI labs have like a pr problem or Speaker 2: something like that. They keep bringing this up, and it's Speaker 2: like maybe they just believe it. Speaker 6: So I think that this concern is quite old, and Speaker 6: I think many people have this intuition that they're like, Speaker 6: I can influence the thing by building it. But now Speaker 6: there's this problem that that logic kind of always recommends Speaker 6: that you continue building more advanced technology or like more Speaker 6: advanced AI systems. And now you have this problem where Speaker 6: there's all of these companies and they all say that Speaker 6: they need to build it because if they don't build it, Speaker 6: another company will. And then even if all the and Speaker 6: they could all have doubts about each other's commitment to Speaker 6: safety or to these principles. Famously, the leaders of the Speaker 6: labs really do not get along. They're not friends. It's Speaker 6: not easy for them to kind of sort out the Speaker 6: safety thing among themselves. And then even if all the Speaker 6: USAI labs kind of agreed to do that, they then Speaker 6: have this kind of external bogeyman of China, Right, well, Speaker 6: what will the Chinese companies do? And so there's this Speaker 6: sense in which just like even the concern is real. Speaker 6: I think a lot of people then who are in Speaker 6: the industry have the instinct that they kind of there's Speaker 6: no guiding principle for what they should do on safety Speaker 6: other than to like build leverage for themselves for later. Speaker 6: And I think that is a concerning state of affairs Speaker 6: for AI development to be in globally. You know, obviously Speaker 6: we're trying to do something different by like informing the Speaker 6: public or kind of giving like you know, you could Speaker 6: imagine that this situation would be better if or like Speaker 6: one gap that exists right now in that picture is Speaker 6: that it's the people building the technology who most believe Speaker 6: that it's going to be destabilizing and sort of all encompassing. Speaker 6: Maybe if the public and governments all were on the Speaker 6: same page and believed the same thing, if it were Speaker 6: true that it was headed there, then there would be Speaker 6: kind of like more time for society to figure out Speaker 6: a response from people who are not trying to build Speaker 6: leverage over the technology themselves directly, or you know, control Speaker 6: the technology via some kind of like public action or government. Speaker 3: Can I just ask very quickly since you brought up Speaker 3: China and I don't want to forget to ask this question. Speaker 3: But Quinn doesn't show up on your like main charts. Speaker 3: I think you did a preliminary assessment of it a Speaker 3: while ago, but like, what's the difference between assessing one Speaker 3: of the closed models in America versus one of the Speaker 3: open source models over in China. Speaker 7: I think one thing to say is that the capabilities Speaker 7: are lacking behind We think that they're they're lacking behind it. Speaker 7: I'm not sure that of it. Speaker 3: They just like don't make it onto the chart. Speaker 7: So we do try to prioritize just because MITA has Speaker 7: has limited resources staff time in particular, that the models Speaker 7: that we anticipate being on the frontier and in general, Speaker 7: the Chinese models have been something like, you know, nine Speaker 7: to twelve months let's say, behind the US models. And Speaker 7: I think the gap by time horizon is probably even Speaker 7: larger than the gap by benchmark scores, where there's some Speaker 7: I'm not sure how scientific I can make this, but Speaker 7: there's some cloaqu real sense or something that the Chinese Speaker 7: models are stronger according to benchmark scores than they would Speaker 7: be on you know, truly held out problems in some. Speaker 3: Sense like gaming the benchmark. Is that what that means. Speaker 7: Or I'm not sure you know exactly how that shakes out, Speaker 7: but something spiritually spiritually close to that. I'm not sure Speaker 7: that's true for all Chinese models. I'm sure it's true Speaker 7: for lots of models outside of China, but I think Speaker 7: that's the least more possibility. Speaker 3: I'm very curious when you talk to external actors in Speaker 3: all of this, and I'm going to group them into Speaker 3: I guess policymakers, investors, and the labs themselves, like who Speaker 3: are you interacting the most with at the moment? Speaker 6: I think that in practice we end up interacting a Speaker 6: lot with AI labs because there's some amount of sorting out, Speaker 6: getting access to models, working with them to set new Speaker 6: precedents and things related to third party red teaming and Speaker 6: third party risk assessment. We think of our audience as Speaker 6: being sort of like high context members of the public, Speaker 6: so the kind of like people, you know, who are Speaker 6: maybe like you do, right, people who are kind of. Speaker 3: Like people listening to this podcast. Speaker 6: So people listening to this podcast people with kind of Speaker 6: who have to make important decisions that will be informed Speaker 6: by the pace of AI progress or like the kind Speaker 6: of profile of AI capabilities. Overall, Because we're based in Speaker 6: the Bay Area, I think we like disproportionately end up Speaker 6: interacting with people who are building the technology and like Speaker 6: closer to it. Partially, I think back to Joe's point before, Speaker 6: I think this is kind of because it is the Speaker 6: case that to kind of care about a lot of Speaker 6: these frontier problems, you're kind of selecting for people who Speaker 6: are building the technology themselves. There's some sense in which, Speaker 6: like the companies in the industry spends more time thinking Speaker 6: today about frontier capabilities assessment than the government does. I Speaker 6: think like one day you could imagine us getting to Speaker 6: the point where the government is like very focused on Speaker 6: this and dedicating a lot of resources to it, and Speaker 6: at that point I would expect Meeter to be spending Speaker 6: more time talking to governments. Speaker 3: That's kind of what I was getting at because our Speaker 3: senses and a lot of the conversations, like we talk Speaker 3: to people and they'll say something about like, oh, it's Speaker 3: important to have a social safety net for an AI Speaker 3: enabled future, but no one seems to be really thinking Speaker 3: about it in a lot of detail. Speaker 2: And when you say, you know, it's easy to imagine Speaker 2: or maybe the government will care more about this, not Speaker 2: so easy for me to imagine. It seems like they Speaker 2: mostly care about you know, data centers and like where Speaker 2: they located and stuff like that. It would be nice Speaker 2: if we had policymakers really looking at like frontier capabilities Speaker 2: and stuff. Still seems kind of a way off, but Speaker 2: it is interesting. You know, you're like talking about like Speaker 2: the sort of like capitalist dynamic, right, there's competition, and Speaker 2: it's like you have a lot of people that are Speaker 2: really worried about, oh, what if the other guys get Speaker 2: to ASI or AGI first, or what if the Chinese, Speaker 2: et cetera. How much does the fact of like free Speaker 2: market capitalism and the demand you know, the big investors Speaker 2: at the VC funds, like they want to return, they Speaker 2: want an ipo if we might get some big AI Speaker 2: IPOs this year in fact, how much do you find Speaker 2: that to be perhaps intention with the safety element? Speaker 6: Yeah, I maybe, Yeah, people on our team wou have Speaker 6: different views on this. I personally don't feel there's, yeah, Speaker 6: there's some thing you're like investors are key decision makers. Speaker 5: And you know they're people too. Speaker 6: That sounds strange to say investors or people do I Speaker 6: sound like Mitt Romney or something. But I think that, like, Speaker 6: I think that the element of this that feels like Speaker 6: it could be intention is if you build a bunch Speaker 6: of financial obligations to keep kind of the pedal to Speaker 6: the metal no matter what the risks are going into Speaker 6: the future. So, like, one thing I think a lot Speaker 6: about is if you're like building up a huge amount Speaker 6: of debt to build data centers and then say that Speaker 6: you do find evidence that you're now worried about about Speaker 6: the you know, loss of control from AI systems, you Speaker 6: do find instances of AI systems going rogue. Do you Speaker 6: now have like a financial commitment to build up those Speaker 6: data centers and like continue kind of the pace of progress. Speaker 6: I think that is one place where I feel the Speaker 6: tension pretty acutely, Like you're building these expectations into the Speaker 6: market that could kind of force you to continue development Speaker 6: when you otherwise would rather invest more in safety or Yeah, Speaker 6: like it at least gives you a kind of financial Speaker 6: obligation to continue scaling at least compute. I think that Speaker 6: like the people themselves being informed about the progress does Speaker 6: not seem bad to me. I think it's like good Speaker 6: in some ways for everyone to be on the same Speaker 6: page about capabilities that could be related to subverting human Speaker 6: control later on. But I think in the world beyond Speaker 6: like the information that Meter shares, I do think there Speaker 6: is a tension, like the fact that private companies are Speaker 6: building this I think could cause really acute tensions in Speaker 6: the future where people make these commitments that they wouldn't Speaker 6: if they were trying to like slow or you know, Speaker 6: maximize social resilience of the technology. Speaker 5: Yeah. Speaker 7: I'm not sure how these things shake out, but I Speaker 7: think there are some forces on the other side, right, Yeah, Speaker 7: you know, some safety promoting technologies quote unquotes or techniques Speaker 7: do make the models more useful, you know, if they're Speaker 7: better complying better complying with your whale in some sense, Speaker 7: and so have capitalist incentives standard capitalist incentives to invest Speaker 7: in that kind of research. Maybe that doesn't cover you know, Speaker 7: the broad suite of safety research that seems important. It Speaker 7: certainly doesn't rule out capabilities progress as being an important Speaker 7: taxis on which you do want to scale. But you know, Speaker 7: I think there are some some forces in each direction. Speaker 3: Since you mentioned compute just then, can you talk a Speaker 3: little bit more about I guess the relationship between like Speaker 3: the time horizon improvements and the cost of compute at Speaker 3: the moment, and like what you've actually seen and how Speaker 3: that impacts it. Speaker 7: Yeah, so, so one extraordinary fact from my perspective. I'm Speaker 7: not sure how to how to fit these facts together, Speaker 7: but something like the R and D spend on compute Speaker 7: of these companies has risen exponentially, of course, and in Speaker 7: fact it's risen exponentially at essentially the same rate as Speaker 7: time horizon progress. You know, I think there's nothing necessary Speaker 7: about that. You know, it doesn't mean by itself that Speaker 7: if computer progress lows then capabilities progress will also slow, Speaker 7: but you know, it's clearly an important input into into Speaker 7: AI progress. I expect that to continue to be through Speaker 7: in future. Sometimes people ask us if we think it's plausible, Speaker 7: or how plausible we think it is that that capabilities progress, Speaker 7: this exponential capabilities progress might slow down at some point Speaker 7: at some point in the future. And you know, one Speaker 7: reason it seems it's hard for me to consider it Speaker 7: plausible that it will slow down in the next at Speaker 7: least small number of years. Is that a lot of Speaker 7: those computes are and the investments basically already bigged in, right, Speaker 7: Like the data centers have already been built, you know, Speaker 7: plans for data centers even beyond twenty twenty seven twenty Speaker 7: twenty eight are presumably you know, coming coming to fruition Speaker 7: coming about, and so some of these input investments are Speaker 7: already baked in in some sense. So it would be Speaker 7: surprising to see capabilities slow to the extent that computes Speaker 7: has been has been an important input. After that, maybe Speaker 7: maybe you need to think about, you know, other arguments Speaker 7: for how capabilities might slow. Speaker 5: But that's roughly how I think about it. Speaker 2: There's a very good or interesting critical subject post called Speaker 2: against the Muter grav by someone named Nathan Woodgen who Speaker 2: brings up one an interesting point that I wouldn't have Speaker 2: thought of heading out Reddit, which is you're paying the Speaker 2: software engineers to come in and perform these tasks, right Speaker 2: it seems, you know, maybe this will be the last Speaker 2: job of humans, is just doing benchmark If I were Speaker 2: like a good software engineer and you say, Joe, come Speaker 2: in and do this task. Speaker 4: How do you prevent me? Oh man, this is taking Speaker 4: me a long time. Speaker 2: Mean, why I keep getting one hundred dollars an hour Speaker 2: for like looking at my computer and time? Who this Speaker 2: is tough. I'm gonna have to come back tomorrow and Speaker 2: keep working on this. How do you avoid the sort Speaker 2: of conflict of interest where the person who's paid to Speaker 2: work on this problem may be encouraged to take as Speaker 2: long as possible to solve it, and with only three Speaker 2: people working on it at times, I don't know, like Speaker 2: this does not It seems like a conflict of interest Speaker 2: to me. Speaker 7: Yeah, So the shulds onset is, you know, in general, Speaker 7: we are incentivizing these people to complete the task because Speaker 7: you know, it's possible, in particular, to complete the task Speaker 7: faster than that is who are attempting the same task Speaker 7: the time that it would take for them. Speaker 3: They task a bonus if they do it faster than Yeah. Speaker 7: Yeah, approximately, there's a bonus if they complete it faster Speaker 7: faster than anyone else. Speaker 5: You know. Another thing to say. Speaker 7: Is I think it just is true that our baselining methodology, Speaker 7: or the ways in which we compare to humans in Speaker 7: some ways leaves a lot to be desired that you know, Speaker 7: ideally we would have invested, you know, one hundred times Speaker 7: as many resources in having one hundred baselines human basedlines Speaker 7: per task, and those would have come from, you know, Speaker 7: perhaps the very best software engineers or machine learning engineers Speaker 7: in the world. Maybe that would be the Maybe that Speaker 7: would be the comparison that we're making. And indeed, we'd Speaker 7: be doing all of this procedure over many more tasks, Speaker 7: not just many more tasks, many more tasks, over wider Speaker 7: task distributions than just software engineering or machine learning engineering. Speaker 7: I mean, I do think time horizon still represents progress Speaker 7: over over what's come before in the science of measuring Speaker 7: AI capabilities. But you know, in some ways I'm sympathetic Speaker 7: to a lot of criticisms of time horizon. I do Speaker 7: think that some of the details, at least for the Speaker 7: work we've done so far, you know, aren't going to Speaker 7: matter as much as you might naively think. So trueing Speaker 7: the shortest baseline time that we end up observing or Speaker 7: the longest time you know, it's actually not going to Speaker 7: make that much difference to the final measurements. Speaker 5: You know. Speaker 7: Of course, we do think these people are talented software Speaker 7: engineers or cybersecurity people or someone depending on the task. Speaker 7: But you know, perhaps we could have found even more Speaker 7: talented people. They would have completed it in half the time. Speaker 7: And so you know, naively, it would seem like the Speaker 7: time horizon that we estimate of these models would be Speaker 7: half as long as we actually end up observing. But Speaker 7: of course that that wouldn't change the doubling time. It Speaker 7: would mean you'd get to the same level after another Speaker 7: four months. In some sense, the big picture that I Speaker 7: want time horizon to point to is less this like Speaker 7: Opus four point six is twelve hours in particular, and Speaker 7: more that we're seeing this remarkable pace of progress that Speaker 7: shows no signs of slowing in the recent past, and Speaker 7: I think in the near future as well. You know, Speaker 7: in fact, it shows some signs of speeding up. Speaker 3: Well, I was going to ask about this because I Speaker 3: think recently the statistic that you would always hear was Speaker 3: like a doubling every seven months something like that. How Speaker 3: fast do you see it going in the near future? Speaker 7: Yeah, so I was a doubling over every seven months. Speaker 7: Person that there was there was controversy in our team Speaker 7: about about what to believe here because when we originally Speaker 7: published this work approximately a year ago, you'd see, you know, Speaker 7: if you plotted a single straight line, a single exponential Speaker 7: you'd get something like, you know, six or seven months, Speaker 7: let's say. Speaker 5: But if you. Speaker 7: Restricted to just the time since I think JPT four Speaker 7: to oh, since the twenty twenty four models onwards, you'd Speaker 7: see something closer to this sort sort. Speaker 5: Of like four or five month trend. Speaker 7: And some people believed in that, and you know, some Speaker 7: people like me had the intuition that, well, we have Speaker 7: so few data points, we should we should really be Speaker 7: estimating over this larger number of data points than a Speaker 7: large number of. Speaker 5: Data points says every six or seven months. Speaker 7: There are a couple of things that have changed my Speaker 7: mind and made me realize my colleagues were right. Since Speaker 7: since then, One is that for the models that have Speaker 7: that have come out, since you know, what trends has Speaker 7: has better predicted how performance those models would be. And Speaker 7: it's very clear that the answer to that is the Speaker 7: four month doubling time and not this seven month doubling time. Speaker 7: You know that there's some some possibility that could speed Speaker 7: up again. We've seen it. We've seen it speed up Speaker 7: once I think there are some reasons in principle why Speaker 7: you might expect it to speed up again. I think Speaker 7: there are some caveats about this, you know, these are Speaker 7: these are maybe some some takes that my colleagues would Speaker 7: agree with, and so you know, maybe maybe you should Speaker 7: discard that, or you know, you should think that they're Speaker 7: going to commits me in the way that they did Speaker 7: with the with the four month versus seven month doubling times. Speaker 7: I have some suspicion that the tasks that meter is Speaker 7: measuring performance on are you know, in some sense more Speaker 7: and more narrow slice of possible tasks, and in particular, Speaker 7: and more and more narrow slice that is perhaps similar Speaker 7: to the kinds of tasks that you'd expect these major Speaker 7: AI companies to be training on in the first instance. Speaker 7: And so in some sense, we're increasingly more so than Speaker 7: was the case before, measuring progress on the exact types Speaker 7: of tasks that they're trying to get better at. You know, Speaker 7: you might think, for instance, the kinds of tasks that Speaker 7: would make for good reinforcement learning environments, the kinds of Speaker 7: tasks that you can score quickly and cheaply and automatically. Speaker 7: I think that progress is real. I think that progress Speaker 7: generalizes to some extent to other types of tasks I see. Speaker 7: I think we're saying, you know, remarkable progress and these Speaker 7: more messy tsks. For example. Speaker 2: I have one last question, which is like how big Speaker 2: is your team funding? And like also how many people Speaker 2: Meter are basically like really rich from AI and they're like, Speaker 2: you know what, I'm good. I don't need to pursue Speaker 2: like stick around for the IPO or whatever. I'm set Speaker 2: and now I want to work on something that like humanity. Speaker 3: No. Speaker 2: I've seen like there are other independent air researchers and Speaker 2: they talk about this. It's like, I want to be Speaker 2: able to talk about what I saw. Miles Brundage, someone Speaker 2: who has like a little think tank, He's talked about this. Speaker 2: What's like, how many people are like rich already and Speaker 2: they're like, Okay, now I want to work for something Speaker 2: that's public facing. Speaker 6: Yeah, so Meter right now is about thirty people that Speaker 6: we're growing and hoping to grow fast. We are hiring Speaker 6: I should say meter dot org slash careers and yeah, Speaker 6: you were touching before and kind of the thing about Speaker 6: is it difficult to be a nonprofit? Speaker 5: You know, we can't pay people in equity. Speaker 4: We got to get an io, right. Speaker 6: Yeah, there's no no ibo or for Meter, but we Speaker 6: do try to pay competitively on cash compensation, right, So Speaker 6: that's an area where we feel we can like somewhat Speaker 6: compete with labs. And it's true that I think a Speaker 6: lot of our team is just motivated by trying to Speaker 6: kind of do something different like not you know, all Speaker 6: the companies to some extent or in this business of Speaker 6: kind of like building somewhat redundant products kind of competing Speaker 6: for the same role in the world. And Meter is Speaker 6: in a really unique position at the moment where I Speaker 6: think that we have like access and the ability to Speaker 6: communicate these ideas and explain the state of AI research Speaker 6: to a number, like a lot of audiences that might Speaker 6: be hard for like individual researchers inside of a company, Speaker 6: Like we get to talk to a lot of governments directly. Speaker 6: We get to come here and talk with you all, Speaker 6: And that's kind of different. I think if you look Speaker 6: at all the actors that are working on the frontier Speaker 6: of AI research or AI safety, you kind of if Speaker 6: you compare us to AI lab staff, I think that Speaker 6: our work gets to be we get to kind of Speaker 6: every day work on whatever research we think will be Speaker 6: most informative to the like public decision. Speaker 2: Do you have ex AI, not XAI, but ex as Speaker 2: a former AI lab staff who maybe there was a Speaker 2: tender at some point and now they work at mater. Speaker 6: Yeah, we do, okay of those. Yeah, so we do Speaker 6: have some people who previously worked at AI labs. I Speaker 6: do think that as time goes on, I think one Speaker 6: hope that I have is that more, you know, there Speaker 6: will be more and more researchers who have kind of Speaker 6: like made the money that they need from working in Speaker 6: the industry and now are excited and kind of like Speaker 6: lifting all boats by working on kind of like inside Speaker 6: of an organization where the north star can be what Speaker 6: is most informative to the rest of the world outside Speaker 6: of these like relatively small set of companies. Speaker 7: Chris is very polite. I think that's I think that's wonderful. Speaker 7: I'm tempted to be a little bit, a little bit Speaker 7: more aggressive in this conversation. I think we have spoken Speaker 7: through mister's work on some of the most important problems Speaker 7: in the world, problems that are going to define the Speaker 7: future I think for not just the next years, but Speaker 7: you know, coming coming decades, maybe maybe even coming centuries. Speaker 7: And we've also spoken about some of the ways in Speaker 7: which me to work is not might not what you Speaker 7: might want it to be. That there's a long way Speaker 7: to go in the science of evaluating these ais. Why Speaker 7: have we not made more progress? You know, maybe maybe Speaker 7: a couple of reasons. I think clearly the central reason Speaker 7: is that we are bottlenecked on technical talent, on incredibly Speaker 7: capable people to come work on these questions. I was Speaker 7: on a meter work retreat recently where we were brainstorming, Speaker 7: you know, twenty thirty of these what seemed like world Speaker 7: important problems, problems that we think no one else is Speaker 7: going to get to if we do not get to them, Speaker 7: and we are able to conduct research on how many Speaker 7: of those problems, I think it's one. Speaker 5: Two. Speaker 7: You know, maybe if we do an extraordinary job this quarter, Speaker 7: it might be three. As Chris alludes to, I think Speaker 7: if you're interested in, you know, less working on redundant Speaker 7: products at these major area companies and more advancing our Speaker 7: understanding on some of the most important questions in the Speaker 7: world that are going to shake the world for years Speaker 7: to come. Meters is a great place to go. Speaker 6: Well, yeah, One more thing to say about that is Speaker 6: like the vibe inside of Meter is a state of triage, right, Speaker 6: And I think people often tell themselves externally. People might guess, oh, Speaker 6: you know, meters A, it's outside of any of the Speaker 6: AI labs. So the thing it might most struggle with Speaker 6: is things like access to AI models. You know, you Speaker 6: can't do the research you want because you don't have Speaker 6: you're not building the thing yourself in practice, or that's Speaker 6: the story that people always tell us. You have to Speaker 6: build you know, the future to shape it in practice. Speaker 6: I think our experience at METER is that, like when Speaker 6: we want to try new types of research that would Speaker 6: require new kinds of structured access, our experience at this Speaker 6: point has been that AI labs are like pretty game Speaker 6: to play ball on that. And the thing that is Speaker 6: more happening is that we're having to turn down opportunities Speaker 6: to do stuff like that because we don't have the Speaker 6: staff that we need to make those things happen. Speaker 2: Interesting Joel and Chris, thank you so much for coming Speaker 2: on odd Laws. Absolutely fascinating conversation and I appreciate your Speaker 2: taking your time. Speaker 5: Great to have you in the studio. Speaker 6: Yeah, thank you so much, so much, having us. Speaker 2: That was a really interesting conversation to that we're starting Speaker 2: from the end sort of the idea of like, Okay, Speaker 2: here are some really important questions, like let's just set Speaker 2: everything aside. Speaker 3: And there's thirty people working on there, there's. Speaker 2: You know, and like how many people want to do it, Speaker 2: and it's like, okay, we try to match cash comp Speaker 2: et cetera. Yeah, that seems like kind of a tricky Speaker 2: issue if like, if you accept the premise that these Speaker 2: are some big questions we have to get right and Speaker 2: you got to land this plane hopefully, Like that's a Speaker 2: bit of an issue. Speaker 3: Yeah. The other thing I thought was really interesting was Speaker 3: the Chinese models not really making it on the charts Speaker 3: even though, like we know, in the market itself, like Speaker 3: when deep Seak, when that new version came out, that Speaker 3: was like this huge thing where everyone started to panic Speaker 3: and to not see it even like land on the Speaker 3: time horizon chart. It's kind of interesting. Speaker 4: I guess it's interesting. Speaker 2: I mean, I guess I buy the reasoning from their Speaker 2: perspective that the only interesting question from meters perspective is Speaker 2: like the most cutting edge slightly adjacent to the most Speaker 2: interesting chart for like business, right, So it's like, Okay, Speaker 2: we know the deep sea and Quinn and Kimmy and Speaker 2: all those are like very impressive. Do they push like Speaker 2: the very frontier? Perhaps not, but just in general, I Speaker 2: find this space so weird because it's like, here you Speaker 2: have these people who are like clearly quite alarmed at Speaker 2: the potential here, and most people, I think, look at Speaker 2: these charts and they say like, wow, this is like Speaker 2: I want to invest in this, or this is. Speaker 4: Like no, I know, I know. Speaker 3: Like that's why my first question was like, you're here Speaker 3: for AI safety purposes, but everyone seems to get excited Speaker 3: about the line go up charts right, Like there's a Speaker 3: disconnect all connected. Like I say, when an industry basically Speaker 3: says it's worried by itself, you should pay attention. Speaker 2: It's really strange. This gets back to, you know, very Speaker 2: It's very strange where you have the CEOs of these Speaker 2: companies who are in many cases the most alarmist, and Speaker 2: there's this sort of cynical thing. And I don't totally Speaker 2: discount the cynical interpretations like oh, they're saying this because Speaker 2: they want to get investors and so forth, and they Speaker 2: need all this money. But look, it was also true Speaker 2: that open AI and Anthropic but open AY a little Speaker 2: more were like founded with these very exotic corporate structures Speaker 2: of like a private company owned by nonprofit et cetera, Speaker 2: which they presumably did because they took pretty seriously the Speaker 2: fact that this technology is science. It was like very Speaker 2: strange and not just like it's not just enterprise office right, Like. Speaker 3: They were self limiting in a way. Speaker 2: One other interesting thing too, that this idea is like, okay, like, Speaker 2: first of all, what's the difference between seven months and Speaker 2: four month time doubling? Speaker 5: Not much? Speaker 1: You know. Speaker 3: It's like these people's like, oh, I can't but it's exponential, Speaker 3: isn't it. Speaker 4: I guess it's exponential, But it's still funny to me. Speaker 2: It's like, oh, I think like AI is going to Speaker 2: destroy all white collar work in two years, and someone Speaker 2: else is like, no, no, I think it's gonna be three years. Speaker 2: Is if that makes any different whatsoever? But one thing Speaker 2: to consider all sort of alluded to this. You know, Speaker 2: you had like open ay shut down. It's like video efforts, Speaker 2: et cetera. So perhaps part of the story is just Speaker 2: this intense focus now on the software engineering side, as Speaker 2: what these labs are working in Yeah, and sort of Speaker 2: like all these other side quests are not as important, Speaker 2: So maybe we will see even more rapid progress on Speaker 2: some of these technical benchmarks, because clearly, from the labs perspective, Speaker 2: that's where the action is more than some of these Speaker 2: consumer things like making making images or videos. Speaker 3: Yep, all right, shall we leave it there, Let's leave Speaker 3: it there. Okay, this has been another episode of the Speaker 3: Auth Thoughts podcast. I'm Tracy Alloway. You can follow me Speaker 3: at Tracy Alloway. Speaker 2: And I'm Joe Wisenthal. You can follow me at the Stalwart. Speaker 2: Follow our guest Chris Painter He's at Chris Painter yup. Speaker 2: And Joel Becker He's at Joel Underscore b k R. Speaker 2: Follow our producers Carmen Rodriguez at Carmen armand dash Ol Speaker 2: Bennett at Dashbot, kil Brooks at Kilbrooks and Kevin Lozano Speaker 2: at Kevin Lloyd Lozano. And for more odd Laws content, Speaker 2: go to Bloomberg dot com slash odd Lots where the Speaker 2: daily newsletter and all of our episodes and you can Speaker 2: chat about all these topics twenty four to seven in Speaker 2: our discord Discord dot gg slash lots. Speaker 3: And if you enjoy Odd Lots. If you like these Speaker 3: AI episodes, then please leave us a positive review on Speaker 3: your favorite podcast platform. And remember, if you are a Speaker 3: Bloomberg subscriber, you can listen to all of our episodes Speaker 3: absolutely ad free. All you need to do is find Speaker 3: the Bloomberg channel on Apple Podcasts and follow the instructions there. Speaker 3: Thanks for listening.