
As AI is reshaping the tech world, so it in turn waits for something greater to come along and reshape it, and that unstoppable force might just be TypeSafe AI's new model, Jev. You can listen to the latest episode of The Kettle on this page, Spotify, Apple Podcasts, or YouTube. You can also subscribe on those platforms to be notified when a new episode goes live. Jev's not chatty, it doesn't want to tell you a story, and it won't really do much besides answering multiple choice, ranking, and binary yes/no questions to a degree of probability, but in lots of cases that's probably all you need from an AI. Jev is also dirt cheap to run and ultra fast, and has proven to have a lot of use cases in just a week on the scene. This week on The Kettle, join host Brandon Vigliarolo, senior reporter Tom Claburn, and contributor Joab Jackson to chat about how Jev is reshaping the AI landscape and what it might do to the future of the AI industry. A lightly-edited transcript is below: Brandon (00:01) Hey folks, Brandon Vigliarolo here with the latest episode of The Register's Kettle Podcast. This week, we've got what seems like an AI revolution on our hands. and no, you have not wandered back in time to 2022. We're talking about TypeSafe AI's Jev, which has gone viral in the AI world by promising to do some real work for far cheaper than its large language model cousins. And with me to discuss this seemingly new paradigm in AI, our senior reporter Tom Claburn. And contributor Joab Jackson. Thanks to both of you for joining me. Tom (00:33) Yeah, thank you. Joab (00:34) Thank you. Brandon (00:35) So Tom, you wrote the initial piece on Jev's release last week, and it's still kind of ramped up in its popularity and it keeps getting talked about more and more. So I figured this was a good topic for this week's episode. So explain what exactly Jev is, for those who are not entirely familiar with what it does and how it differs from a large language model. Tom (00:52) Yeah, so Jev is what's called a system one model. It's from a company called TypeSafe AI. And I think the reason that it's gotten a lot of attention is that it has some very credible technical people behind it. I think one of the people involved was with OpenAI. And the basic idea is that the model is highly optimized for doing type safe choices. And so rather than being sort of something that's all natural language and you have to then handle all the ambiguities and of all that, it's made for doing things like making a decision sort of between a set number of options or true / false statements. And it uses a different approach to machine learning. It's called reinforcement learning for calibrated decisions, as opposed to reinforcement learning for human feedback, which is one of the methods used for generative AI. And the idea is I think it's just generated a lot of excitement, partly because there hasn't been a lot of sort of high-profile novel research. There's a lot of people doing work all the time in this field and they're coming up with different ways to do things the Chinese labs are doing a lot of interesting innovations but they don't quite have the same profile because of the people involved, and I think for the US press this was an interesting approach. And it scratches an itch that I think a lot of companies have, which is a way to implement things that need deterministic outcomes. And, 'cause you can't just put a large language model on, I don't know, like an invoice billing system and have it come back with some weird hallucinated answer. You need to know: should this invoice be paid or not? Brandon (02:44) Yeah. And like I know you mentioned hallucinations, right? You mentioned in your story that TypeSafe is billing this as a hallucination-free type of model. I think you described this as not quite being a fair way to describe it, but the idea is that it shouldn't be making stuff up, right? Tom (03:02) Right. I mean it can still make errors. If you give it, say, a list of three options, it'll return A, B, or C for you, but it doesn't pick one absolutely. It gives you probabilities for those choices. Brandon (03:16) Right. Right. Tom (03:17) And so it can still be wrong. Because you've constrained it to three choices, it's not gonna come back with some weird story that's inappropriate for whatever the context is. Brandon (03:30) Right, 'cause it's not giving you really much textual feedback at all, right? It's just, I think there's only three things you can do with it essentially, right? Three types of queries you can put to it. Is that right? Tom (03:40) Right. Yeah, yeah, there's choice, which is like, multiple choice questions. There's scoring them, which is you give it a sort of a scoring rubric and it'll rank things, and then Nool, which is like a true or false thing. But if you were to give it, say, a question about a bill or something and or like where a certain email should go. you could say, well, there's like an 80 percent chance it'll go to a technical department, a seventy percent or 10 percent chance you go to the sales department or whatever. And it would break those down for you. And you would then route it that way. But the reason that it's appealing is that it's very, very cheap compared to other options. If you just let Claude make those decisions, it would spend a lot of time churning and give you some answer that your system could probably use, but it would cost you 50 times as much. Brandon (04:40): And so it's this much faster too, right? 'Cause it's got some different way of processing responses as opposed to generating token by token. It spits all out at once, basically, right? Tom (04:48) Right, right. It handles it all in parallel. So you send out a query and it gets you all the answers back at once, rather than with a traditional language model where it's predicting token by token by token, and because it's a serial operation, it just takes a lot of time. So yeah, it has a lot of possibilities for optimizing things where you need a quick response. Brandon (05:13) And I understand you don't pay for your outputs either, right? Is that correct? You pay for the tokens that go in but it doesn't cost anything coming out, is that correct? Tom (05:21) Yeah, that's my understanding as well. So because how many tokens you're getting back? You're getting back true or false or you know, really... Brandon (05:30) Yeah. True or false, some statistical rankings and some probabilities that it's right in various ways. Yeah. So way cheaper, way faster, and I, from what I was looking at, you can still query this thing using natural language, right? You still ask your questions in natural language. It's not like you have to program a response into this thing, right? I mean, you gotta give it some constraints in a way that you don't give to large language models, but you're still asking it questions in natural language. Tom (05:56) Right. Right. Yeah, Brandon (05:57) ... And then getting getting some probabilities back in response essentially. Tom (06:00) One of the things I included in the story, which is, I mean it's fun, but it also kind of explains a little bit about how it can be used, is they set it up to play Doom. And if you give the model all the player states, like the health and position on the board and things like that, it can provide you know, answers fast enough to play the game. So that's an interesting thing. And you can think, extrapolating from that, there are a lot of, possibly useful operations for this kind of thing. And I think that we'll probably see a lot of parallel efforts where people are trying to develop their own version of this approach because it makes a lot of sense to look at places where if you're building a system with natural language processing, there probably are a lot of places where you'd be perfectly fine with regular old deterministic tree-based ABC choices. Brandon (06:52) Mm-hmm. Tom (06:52) I think a lot of people who have engineered their natural language systems haven't looked at the optimization options. It's like how can we, then make this more efficient? And so sticking Jev in would do that. Brandon (07:00) Yeah, yeah. Yeah, so maybe that's maybe that's a kind of a potential future evolution of this kind of model is this kind of structure finds its way into large language models as a way to speed up and make responses cheaper. But actually speaking of use cases, Joab, you wrote a story this week talking about some of the ways developers are using it. Cause I guess someone stood up a website for basically cataloging all the various ways this thing is being used. So what's what have you seen? What are kind of some of the popular ways it's being being applied? Joab (07:34) Most of what we we're seeing are are demos that could be easily fairly easily replicated, I say, by deterministic. But it's interesting just to see the wave of developers figuring out what they could do with it. It can't provide you any answers you don't already know. If you give it four possible choices to choose from, it's not gonna come up with a fifth choice so you're gonna have to have a kind of roadmap going into it. But so the use cases are for these things that are not quite deterministic but still can be mapped to a limited set of outputs. That's why speeding through games was such a popular kind of demo for these things. Whereas, you know, when you're playing Doom or Tetris or some other game - you're going to have a limit enough choices, but you have the background of well: "I need to get to the end. I need to win as many points as possible." Brandon (08:38) Mm-hmm. Joab (08:39) So it works best for those those projects that have predictable outcomes that you just want to speed through. LLMs can do that, but they take a lot more time because they're formulating a response and how to how to interact back with the user. Brandon (09:02) And they might stick,like you said, they might stick that fifth response in there, right? And that could be a complete nothing response that doesn't make any sense to the scenario because LLMs do still no matter how good they get, they still do that on occasion. Joab (09:15) So yeah, I mean so you got a lot of a lot of classification and routing sort of applications coming up. I just read something about someone using it as a safety check for agents of all the agents, all the Unix commands an agent can commit. Here you don't wanna remove directory or you don't wanna remove this file, you know. So you set up the kind of guardrails for agents. It's really good at that sort of thing. And since it's so fast it doesn't considerably slow the system. And so, unlike LLMs, there's a fair amount of kind of rigorous thinking that the developer must go into before working with Jev. You should have as many questions as possible and they should be specific. it's almost the opposite of the journalism questions where you know you want to ask "yes-no" questions and you don't want to ask open-ended questions. So it's interesting to see how developers are using that. I guess one of the most interesting examples is this one startup that can take a picture of a piece of clothing that you're thinking about buying. You're at the thrift store, you're at a shoe store, you know can this go with my 'fit. And it can come up with a picture and also probability of how likely it is it'll work for you. So you have to kind of do a little bit of transposition of what Jev is looking for and what you're looking to accomplish. But it is faster and cheaper than having the LLM do it itself. And it's less probable or it's less likely to lead to hallucinations. Brandon (11:06) Yeah, yeah, which is fantastic. I mean, I wonder, like you mentioned the agentic safety kind of application. I mean, I wonder if something like this could be used as sort of a general purpose guardrail maker or guardrail machine for LLMs. Do either of you have any thoughts to that end? Could this be applied to LLMs to make them safer? Joab (11:30) Absolutely. I keep thinking back to Bill Gates' comment about Netscape when that first appeared and he had said, rather oppourtunistically, that Netscape's not an application, it's a feature. And so I could see a point where LLMs would include this sort of classifier as a feature. Like do you want an expansive mode or do you want a set of answers? And you could probably do that now, but it takes the LLM a lot longer to do that, and it chews up tokens. So maybe it'll be a feature down the road, ChatGPT or something like that. Tom (12:11) Yeah, and I mean it's it's hard to tell sort of in advance if you know Jev is in and of itself capable of sandboxing. I mean I suspect not. I mean I suspect there are probably, depending on how you set it up, there are gonna be ways that an LLM will around go around because the LLMs have such a broad knowledge of things that, unless you have thought of everything they're gonna think of, by the way I there's this command and I can use this socket to connect to this system that I'm not supposed to get to. Brandon (12:40) Yeah. Yeah. If Jev won't let me do this, then I'll find another way to do it, yeah. Tom (12:44) Right. But in combination with something like a micro VM or whatever, it probably would be a pretty effective way of structuring security around LLMs so it doesn't do, you know, Hugging Face style break-ins. Brandon (12:57) Yeah. Now another thing, speaking of Jev kind of blowing up in popularity and everyone treating it as this new, new, big new thing, I mean, there's been a few people online I've seen asking the question, basically, "is this new?" And almost kind of answering it themselves, right? They're like this is the same sort of thing. This isn't new, right? Zero shot natural language inference classifiers, embedding models, cross encoders, all this stuff has existed for a while and it's doing a lot of the same stuff that Jev is doing. I mean, what makes this unique in what you guys have seen? How is this different or new from some of the other kind of decision tree kind of structures that have already existed prior to this? Joab (13:38) The idea's been around for a while. The moat that TypeSafe has is, actually, from my understanding it is built on an LLM. So as a couple people pointed out, how smart this LLM is, is still open to question. I mean you can build it yourself easily enough, but you still gotta do the fine-tuning of making the LLM as smart as possible for the questions. It's not a full LLM approach. Evidently you send it a question, it comes back with a one-shot answer, but it doesn't ruminate over it. So that's where the cost effectiveness and speed comes from. But the devil entirely is in the implementation of the back end model. Tom (14:35) Yeah. Brandon (14:35) Which we don't know anything about at this point, I assume, right? Beause this is a proprietary product. Joab (14:38) It's proprietary, yeah. Tom (14:39) Right. Brandon (14:39) Yeah. Tom (14:40) As often in software it's not necessarily about the technical capabilities, it's about the marketing. And they put together a very coherent API and pitch, just ease of use counts for a lot in terms of adoption. Brandon (14:55) And five dollars in free credits, from which I can tell, goes a long way with this thing. Tom (15:00) Exactly. So if they can basically sell it well, there's no reason it can't see widespread adoption just because everyone feels like they should bet on it. I mean it's kind of like OpenClaw and Pi. OpenClaw is based on the Pi agent and OpenClaw had its moment in the sun and took off. But the technology was sort of available from other agents. So it's kind of just public perception for a while. And then the hype will settle down and people will figure out what applications is this actually good for and what is it maybe not so good for. But yeah, I mean people have been doing this with classifiers for a while, but there's probably a reason why a classifier, maybe it consumes more resources or whatever. I don't know what the cost for running a classifier is. Brandon (15:44) Yeah. I saw a debate on Reddit where some couple people were talking and one guy was saying, Well, this is really a general purpose classifier in a way that others have not been and the guy countered and said, No, that's not the case. There was a bit of a back and forth discussion going on. So I think there's arguments to be made that this is more capable than a classifier, like Joab said, it's built on an LLM itself, so it's got some of those capabilities in a way that maybe a classifier didn't have. So I know we don't know much about the back end of this thing. If you guys had to kind of extrapolate from what we know about how cheap it is to operate.. could we assume that it's running on systems with less overhead? I mean, could this be something that could run on cheaper, more efficient hardware than an LLM? Or is that LLM back end that we assume is there probably gonna still push it to still need large datacenter operations to function? Joab (16:37) One of the creators of of Jev, I hope I'm pronouncing his name correctly, Diogo Almeida, he had a talk in July, just before Jev was released, about the kind of differences he wasn't discussing Jev in particular, but he had said that up until this time all LLMs had been kind of fine-tuned for human interaction, that's ultimately where you get the hallucinations from, I'm trying to please the user rather than come up with the best response. Brandon (17:12) Mm-hmm. Joab (17:12) And his as of yet unreleased technology is an LLM but tuned towards classification. So I'm assuming that there's still gonna be a lot of upfront work defined to the LLM. Brandon (17:30) Mm-hmm. Joab (17:30) Even if there might not be as much happening on the runtime, which accounts for the lack of output token cost. Because it's not doing that that synthesis, it's just going through the stack or whatever you would call it and coming up with the most likely answer. Tom (17:51) It's also interesting from a meta level, like what the token based economics is doing to software architecture, where you're seeing all this sort of pressure to innovate to cut costs out. And the scenario that Anthropic and OpenAI sort of dream about is everyone's gonna be paying a fortune for their highest end models. But what's happening is that everybody is looking for a way, when you maybe we'll run a high end model when we absolutely need it, but everything else will fail over to a local system or something that's more affordable. And pretty much every process is under scrutiny now. We're thinking, hey, we don't have to pay for that. Here's a more efficient way to do it. So you're gonna end up with a lot of changes in the components that are used to build software to squeeze cost out of it because nobody wants to be paying these very high bills and then hitting limits and things like that. It's gonna force people into innovation for different components like Jev, where you can sort of take certain operations and say "we're gonna put this on this cheaper system." Brandon (18:58) I feel like I read in one of the two of the pieces that you guys wrote that it can be used to essentially be a traffic router, determine where these kind of questions need to go. Can Jev answer it itself? Does it need to go to a local model? Does it have to be forked out to Claude or whatever? So you could even see Jev as kind of an AI front end for business purposes. That it is a way to reduce costs. Jev, I'm assuming, since it's gotten so popular so quickly, I can't imagine that TypeSafe is gonna be the only company with this kind of model. Like I imagine it's gonna be a matter of time before other companies or open model developers are gonna be trying to put out more open weighted open source versions of this available from places like Hugging Face that you can run locally, theoretically. I mean I'm guessing that's probably not gonna take too long if it hasn't happened already. Joab (19:46) You can see it also just more broadly with Facebook's Muse, which is another sort of end run around the LLMs. It gives you a select number of options. Do you wanna connect with this program? Do you wanna do this? And so it's taking an LLM based service and carving a smaller, more specific service, more one attuned to the end user than what the LLMs themselves are offering. So not only classifiers, but all sorts of... I wouldn't say niche, but more end-user-specific type AI services on top of the LLM. Just have the LLM do, you know, the expensive thinking part, but you know, offload that as quickly as possible. Brandon (20:34) Right. Yeah, yeah. It's gonna be very interesting to see because I mean this is still new, right? I think it just came out last week. Is that right, Tom? When it launched. So this is new, right? And it's already gaining a lot of traction. So clearly, again, it's another example of how hungry the industry is for AI services that are not going to completely burn through token allotments that aren't gonna cost an exorbitant amount of money. Here we are. It's the start of the Jevolution. Joab (21:04) Also it keeps developers in business. Brandon (21:07) Sure, yeah. Joab (21:08) There's a lot of work that a developer has to do to really make this run correctly. And so it's rather than say, Anthropic, saying "we'll do everything for you," - no, there's still a lot of work that needs to be done at the user level or at the service level. Brandon (21:24) Yeah, so maybe that's another solution for vibe coding garbage too. You can still get some AI-like performance out of this thing without having to resort to letting Claude do all the work for you and then trying to take credit on GitHub or something. Yeah, we will see. It's gonna be an interesting few weeks, few months to watch to see kind of where this settles in, in the larger AI landscape. Cause it's definitely changing things just like AI has been doing for the past few years. And there's definitely gonna be more to come. I mean, we'll see how long it takes for Jev to do something horribly untoward, too, given how easily these AI models are being turned to nefarious ends. Joab (21:58) Ha ha ha ha. Brandon (21:59) No matter what happens and what shape the the Jevolution takes, we will be here to talk about it on the Kettle and we hope to see you there. (R)