TH
The Vergecast
The Verge
The Philosophy of Human Thought
From The one AI detector people actually trust — Jul 16, 2026
The one AI detector people actually trust — Jul 16, 2026 — starts at 0:00
Hello and welome to the Verge Cast, the flagship podcast of homeomegrown human Witing. I'm Jake Kasternakis, exxecutive Vedor of the Verge, and today we're talking about AI detection and the one system that might actually work I have been on the hunt for a reliable AI text detector for a while now And I know I'm not alone Here's a call we recently got from a listener Hey, my name is Adan to AI plagiarism or AI text detectors measuure? Are they reliable? Can they reliably detect AI generated tech Recently there was an article at my college newspaper. And it's one hundred percent AI generated according to FatGBT to the publication and they refused to take it down, stating that AI detectors are not reliable Thank you, Be For the longest time, this has been the refrain. AI detectors aren't reliable So maybe a student's paper or an executive's LinkedIn post looked like AI there wasn't a sure fire way of knowing changing because now the thing I keep hearing is AI detectors aren't very reliable, but Pangram says it might be AI So today, we're talking to Max Spirro, the CEO of Panram, which makes what might be the first trusted AI text detector on the market We're going to talk about how it works. how much we can trust it and what we should do with its findings But first, here's what's happening on the verge today This is ninety seconds on the verge for Thursday, july sixteenth, twenty twenty six One pllus is exiting the U.S and Europe The company made the announcement today twelve years after first making a splash with the One plus one This is a real bummer for smartphone fans. One Plus had its ups and downs, but the company genuinely was a pioneer in low cost high spec devices that could go head to head with the big flagships David Mmell has a great piece in the Vverge today about how the U.S carrier system is a big part of what killed One Plus Its prices may have been great, but they never looked that great besidide an iPhone that only cost four dollars a month on contracts. Next, the EU is forcing Google to make Android open up more in Europe The European Commission said today that competing AI assistants need to get the same level of access as Gemini That means letting them be activated by voice commands and giving them the ability to control apps Google argues that this presents security and privacy risks, but as of now, it's on the hook to make it happen by july twenty twenty seven The EU is also updating rules, requiring Google to share search data with competitors in Europe That now has to include AI chat bots. Finally, do you want a couple companies to be able to dominate US airwaves FCC Chairman Brendan Carr does. He's planning a vote to end the national ownership cap, which currently prevents broadcasters from reaching more than thirty nine percent of U.S households Instead, he wants to be able to review and approve deals that would violate the CP on a case by case basis letting the FCC consider factors such as quote, pointoint diversity Fun fact, the thirty nine percent cap is in fact enshrined in law So this is definitely going to end up in court. Another win from the C FCC You can read more at theVerge. com dot That's ninety seconds on the Verge for Thursday, july sixteenth, twenty twenty six Support for the show comes from serervice Now. AI is moving fast across the enterprise, but without visibility, it's just chaos Different tools, different models, different teams using AI in completely different ways Service Now turns that chaos into control With the AI control tower, you see all your AI across the business in one place. What it's doing, what it's done and what it's about to do So you stay in control putut AI to work for people Visit sererviceNow. com Support for the show comes from Mongo DB AI Assisted and Aentic Codating is helping you build faster than ever But if your data layer is still a bottleneck What's the point Instead of wrestling with rigid schemas or translating data formats MongoDB's native data model mirrors the language LLMs already speak. It ships at the speed of AI, is acid compliant, and scales to handle massive Fortune five hundred workloads Developers have a word for that kind of reliability. Actually, five words. It's a great fricin database. Start building at Mongob dot com slash AI All right, we have Max Spirro, CEO of Panggrram. Max, thanks so much for joining us. I'm really interested in talking about what you guys have been building Yeah, thanks so much for having me. Really excited. So there have been these AI authorship debates for a couple of years now. and earlier this year, I started to notice them playing out a little bit differently. I feel like it was always people saying, AI detectors aren't reliable, they aren't reliable. And then suddenly people were saying, well Pandgram says, this is AI. And this tool had emerged as a name that people actually trusted. My question to you is what happened there? and how did Pan Gram, at least from my perspective so quickly get this reputation as a trustworthy name in AI detection It's funny for me, it doesn't feel at all like like we gained it quickly. So I think We've been around for Almost three years now. that But I think for Um, the first two ish years of our life, we were just like fighting this narrative on like AI detectors don't work. We were publishing papers, publishing technical reports Um and then I think we like slowly were embedded more in the research community And then we had a couple like Big name researchers who decided to benchmark Pangram and publish results. And I think that's when people start to realize like, oh These guys are legit. They're not just like lying about their accuracy. they They're real So did you start out as purely a research product before productizing it? Or at what point did you get the model that kind of cracked it and went, oh, this is reliable enough that we kind of want to brag about it? Yeah, I mean, definitely it was like just research at the start. It's just me and my co founder. we both have like AI and machine learning backgrounds and we're just like, this sounds like a fun problem to try and And it seems like nobody has really cracked it to the degree of accuracy that people need or want And so it took us U Probably I think a bit over a year to get our first version of the model that really had this like flagship False positive rate was significantly lower than everyone else at, you know, early on It was a one in one thousand false positive rate. Sozero point one percent and Now today, it's one in ten thousand, which iszero zero one percent. and I think that's like Confidently, this is low enough that people are able to confidently point at panram results and be like, o, this is what it says So why don't we back up to that? What is the approach that you took that got you to that one in ten thousand false positive rate that you guys are advertising So It's a method of machine learning called active learning. Essentially what happens is we take a model that's okay, it's decent And then we say, scan this really, really large corpus of Q inter written text. and find out which examples we have errors on. what And what this essentially does is it finds examples that are close to the boundary between human and AI And then we take those. And then we say for each of these documents say it's like a a Yelp review on Denny's, Th then we'll ask AI to also create A yelp review about Denny's in the same style. And then so now we have a human side and then an AI synthetic mirror And so we We train on these The human example plus the AI synthetic mirror. And our model is able to learn the difference in stylistic choices that between these two examples. How did you tap into that Like whereere did that idea originate from From like these core machine learning ideas where you want as large of a data set as possible and you want as diverse of a datas set as possible. So I think that's sort of how we landed on synthetic mirrors because otherwise if you ask if I just just ask AI for ten thousand essays I'm going to get like nine thousand essays that sound like very, very similar. So instead what we have to do to diversify our data is to The AI essays mirror A human essay. so that's really interesting. So you're getting this huge body of human work You're then mirroring that with AI versions and training it on the trickiest subset of the ones that your detector can't always tell the difference. Is that right? I'm looking for distinctions between how human writs and N AI rights? Exactly. Yeahah. and training on these hardest examples is a really important part of it as well because otherwise It's just not, there's just not enough signal. like it's For most pieces of text, it's actually very obvious to the detector if it's Um, If it's human or not, Is it riddled with typos? Is it messy or personal. and so by looking for these edge cases where like it maybe plausibly could have been written written by AI We get like much higher signal on what are the actual AI signals So there's a really interesting thing there where Panggram itself is using an AI model to assess AI text. And I've used your service and you have this feature where Pan Gram can kind of flag parts of a body of work that it it believes are signs that tip it off as being something that was written by AI. But at the same time, you guys have this sort of warning or this caveat saying Oh, actually, this isn't what the model is looking at. We don't really know what the model is looking at. Am I understanding that right? So it's the model is sort of a black box, like you guys can't quite Hell what is triggering it to tell What is AI and what is human Yeah, I think usually it's like a more holistic story then the clean story is like, oh, this sentence tips us off that it's AI But the actual answer is that the model's really looking holistically at the document. And there's a whole bunch of micro decisions Each micro decision alone is a decision that AI would have made and a human wouldn't have necessarily always make that decision Um When you like aggregate all of these micro decisions together, you can have high confidence that the document was AI generated. So yes, these Things in our dashboard, we have some like supporting evidence. We've mostly taken these from like the Wikipedia signs of AI writing And we just kind of show these to people as a way to personally train yourself on how to detect AI writing. Is that saying that you Panram, despite being, you know, the flagship AI detector, you don't have your own signs of AI writing. you're you're relying on pedia to kind of turnurn it into something that is like English readable Yeah, in a sense, yes, I think the like English readable side is the hardest thing. We've been doing some really interesting work This field of research is called interpretability T. So So in our case, we are looking at what neurons in the like panram neural network are activating when it sees AI? and how does it fluster AI text differently from human texts. So we had a pretty cool blog post that we put out recently where we found that Even though we're not training the model specifically by saying This AI text is from Claud and this AI text is from Chat GBT. Our model still learns what model family the text is from and is able to Do a pretty good job at Clustream texts from different models separately You know, O one thing I'm curious about too is I've seen people go to ChatBT or, go to Claud and ask them to assess whether writing is AI because those things they are AI. they generate AI. And my impression is that those things have zero capabilities specialized for this whatsoever But there are other specialized AI detectors out there. And those are the ones that I guess have not developed this reputation of reliability. I'm curious Do you have a sense of what they're doing wrong or or not doing that isn't giving them this success rate Yeah. So a lot of the early AI detectors, certainly, we were not. anyywhere close to the first But a lot of the early ones relied on research which said There is this metric called perplexity which might be a good a method for detecting AI text. This is a measure of how surprising a piece of text is to language model. So if you take the sentence like boy ate a bowl of soup. ret low perplexity. every word is U expected Whereas if you have a sentence the boy ate a bowl of spiders, spiders would be a high perplexity word because that's not expected. Right. So If you look at the way AI language models are trained. They're trained to produce low perxity unsurprising. sentences because if it's surprising, it's more likely that it's wrong. So you can build an AI detector that measures perplexity of a text and says, if it's low perplexity It's AI, and if it's hyperplexity, it's human because humans write in a more surprising way. Breaks down, however. in a couple of cases, so any text that is memorized by the AI is going to be low perplexity. For example, like the Declaration of Independence. So if you're wondering why like you put the Declaration of Independence into some random AI detector and it says it's AI, that's why. It's because it's low perplexity. the AI model has memorized it The other issue with it Is English language learners? Also write in simple language which is low proplexity, which then gets flagged as AI by these detectors So that's sort of the problem with these early detectors. And the problem with perplexity in general is it's not really a metric that can be improved upon Like it just It just is what it is. It's like measuring I don't know, like the density of a liquid. it is just a single metric. There's no likeike improving upon it Got it. And so just to break those apart, the old style detector with perplexity They're essentially looking at how how homogeneous a document is whereas Pangrram perhaps to simplify you have trained on the specific patterns in AI text Yeah,re we're sort of like learning the micro decisions that these AI language models make consistently So I guess big question, if I see a pangram result and I see a lot of them these days Should I trust it, right? How should I think about Pangram assessment Yeah, so I mean, I I think Ping Ram iss very accurate, of course, but what most of our benchmarks say is The false positive rate is one in ten thousand. so Um, There's a chance that it could be wrong. Hundreds of thousands of things are scanned by Panram every day. so like there's going to be a few errors. But I think the longer the text is, the more confident we can be that it's correct because we just have more more data So if it's like a fifty word tweet That's flagged by Pengramas AI. like it's most likely correct But I think there's like greater error bars on How much of it was AI? was it actually just AI assisted? Whereas if we're looking at like a on eighty thousand word novel and Panram says this thing is ninety percent AI. like we're very confident that it's at least ninety percent. That's a majority AI written I've noticed your tool is able to do that where it will give some it'll tell you how confident it is in a result. U I was messing around with it and kind of interspersing human written text and AI written text and Um There was one big chunk that was kind of fifty, fifty, and Pan Gram told me There wass like I It low confidence human written. U Whereas there are other times I've seen it tells me you know, it thinks that something is partly human and partly AI. So how are you coming to that a decision when you are you are deciding, oh, I see bits of both in here Yeah, so a lot of what we're doing is we're scanning the document first off holistically, but also like in parts and we're saying like Let's look at this part and does this part look like AI or human or assisted And then for each part we kind of like have a score. and then we are gonna like staple these together and aggregate them and say like, well, we looked at twenty parts of this document and like of them look like they're AI. so we can say it's about like fifty percent AI Support for the show comes from even realities When you walk into a presentation, you have some options to help you remember your notes. You can bring a tablet Wprite all of your talking points in ink on the back of your hand or try your best to simply memorize them Here's a fourth, much more innovative option. You can wear them on your face with even realities. Even G two are productivity smart classes designed to keep real time support right in view Teleprompting, conversation support AI assistantance, and more, they help you stay on top of work and daily life And unlike most smart glasses, they're designed to look and feel like premium eyear With no camera and a lightweight thirty six gram design you can wear all day The more context you give them, the smarter they get, adapting to how you work and what you need To learn more about E G two Go to evenrealities. com and see how everyday smart classes keep helpful information in sight. So you can stay productive and hands free throughout the day And for our listeners, use promo code Verge at evenrealities d. com to get ten percent off E Ring one and or even clip when you add them to your EG two order That's even realalities d. com Promo code Vverge ort for the show comes from even realities For a long time, the reality of smart glasses was that they were these bulky, inelegant pieces of tech that were much cooler in concept than in practice, a far cry from the sleek, stylish sucessory of our sci fi dreams that actually provides real functionality Well, even realities has that fixed Even GCU are productivity smart glasses designed to keep real time support right in view with teleprompting, conversation support, real time translation AI assistants, and more They help you stay on top of work in daily life And unlike most spark glasses, they're designed to look and feel like premium eyear with no camera and a lightweight thirty six gram design you can wear all day The more context you give them, the smarter they get Adaptting into how you work and what you need To learn more about E G two, Go to even realalities. com and see how everyday smart classes keep helpful information in sight so you can stay productive and hands free throughout the day And for our listeners, use promo code Verge at evenrealities d. com to get ten percent off E ring one and or even clip when you add them to your even GCu order That's even realalities d. com promo code Verge Support for the show comes from Mongo DB. AI assisted and agentic coodating can help you build faster than ever. But if your data layer is a bottleneck, what's the point Mongo DB actually gets out of your way MongoDB is a unified AI ready data platform that empowers you to build scalable, generative AI applications It eliminates the need for separate, specialized vector databases by combining a flexible document model natively with semantic vector search, full text search, and real time operational data Instead of wrestling with rigid schemas or translating data formats, MongoDB's native data model mirrors the language LLMs already speak. Plus, Mongu DB gives you the flexibility to ship at the speed of AI. The acid compliance lets you sleep soundly at night and scales to handle massive Fortune five hundred workloads. Developers have a word for that kind of reliability Actually, five words It's a great freaking database Start building at MongoDV d. com slash Ai I'm Nil A Pppellll, editor and Chief of Virge and Coder is my show about big ideas and other problems. Today we've got the first of a two part series on the systems that run the world I'm talking with Bark Butler, the CTO of Proton a company that makes private and secure productivity software It's impossible to create a back door that can only be used by the good guys No company is going to go to jail for you Often the response is, well, if you change the legal foundation here, we will leave. Yeah. How real is that It's dead serious. With all due respect to Swiss authorities and everybody else, I think it would be suicidal. to continue down this path. subscribe where your podcast. This series is presented by Comcast Business A big scandal recently around, I think AI detection was there was this Commonwealth prize. they awarded big short story prize to a piece called The Serpent in the Grove. A lot of critics thought it read like AI. Panram assesses it as one hundred percent AI generated. But the author, Hamir Nazir, he insists he wrote it himself. He just was interviewed by the Atlantic. He said, oh, you know, I used voice to text. to actually write this story and maybe that created some linguistic oddities. I'm curious how you think about that I am very skeptical of his claims I First off, I think like just my personal intuition, it just reads like total AI, like it reads like, you ask Chat GPT to write a literary prize winning essay and like this is what it would say, which is you know, kind of vapid and has a bunch of empty metaphors Um and is like fake pretend deep At least that's like my I read on it. I really don't want to be like too critical of the guy, but I think there were also like a lot of inconsistencies in Interview which also make me suspicious. For example, like he was asked what his favorite author was. He mentioned a couple authors and then the interviewer asks, hey, okay, what's your favorite piece from this author? And then he's like I can't actually remember piece from the sal, which is is like very Um I think that lights up some like some alarms for me of like, is this person really just like B sitting or did he actually spend a lot of time on it Also, I think like the speech to text is kind of surprising to me. like I I like don't see how that would trigger pangrram. unless you took speech to text Um, like rambled for five minutes or whatever and then asked AI to turn that into a coherent essay. That sort of gets to something interesting where Part of the evidence here is first people read it and people who are really familiar with AI writing said, I'm noticing some things here Th thenen people ran through Panram and Panagram said, yeah, we think the say I. And then people asked him and kind of assessed the evidence that he provided and some of someome parties You know, I think including the prize board found it compelling. I think other parties pererhaps I don't know that the Atlantic issued judgment, but certainly suggested some skepticism. as part of that interview, which I think is pretty well warranted You know, it's interesting. Granta, which ran the story said that it asked Claude to assess whether the story is AI. And I believe Claude said that it wasn't. And I think that's a very reasonable approach if you listen to all the leaders of these AI giants who are saying they've built super intelligence, but we both know that this is, you know, this is a specialized skill to detect AI writing, and that check was in all likelihood, like probably basically useless. I'm curious You run this program that is supposed to be reliable. Who is using Pangrram right now A lot of industries seem to be completely unprepared or unaware at this point of how to assess AI writing. I think it's really only been in the last six months where I think the tides have really turned against AI writing. I think for a while people were holding out and saying, well, you know, maybe like AI writing is going to become more common and is going to be like generally accepted. And now we're realizing like, oh, there's actually a lot of use cases beyond our initial ones. So who's using Panram? I think like There's a very big cohort of educators and schools and universities who use Pangram to F a student is cheating on an assignment if they're using AI to ully generate their essays, etcet I think we also have users in publishing, either they have like a magazine or they're an agent or something like that and they want to use Panram to just make sure that everything they are going to publish is above board. We also have like AI companies who use Panram to make sure their data is clean, for example. you don't want to pay an expert for Um to like write up some data or like write an explanation or solve a problem. But actually instead of doing that, they're just feeding it into Chat GPT and then sending it back to you. So I think that's been a growing business as well. That is a wild problem of their own creation. That's true. I'm curious. what do these deals look like for you? Do you have a specific you know, product for these companies or for educators or is it just Come on in and buy a bunch of pan Gam credits and scan away Yeah, we try to bring Pangram to where people are at So like with high red we They use a learning management system called Canvas typically. And so this is where students will submit assignments and receive grades. And so we integrate directly into Canvas automatically score u assignments with Pangrram and then just show that to the instructor And so kind of similarly, we work within people's like content management systems Wherever they work, we're trying to bring Panggram there. If I'm a professor, what should I do if I get a positive result? Do you think that that's enough to fail a student The first step is always talk to the student. I think the reason that somebody would use AI It It's like a symptom of a deeper problem It's either the student didn't have enough time to complete the assignment They didn't feel like they had the understanding to complete the assignment or they like don't care about your class enough. And so I think all three of these are like something that you probably want to dig into deeper rather than just like giving a zero. That a very polite way of saying, though you think the assessment is correct and they should chase that down I mean, I see this all the time where like Educators, they get better results if they do something beyond giving a zero and they suspect AI use because I think that doesn't really like solve the core of the problem, which is that this student feels incapable of like producing an assignment Well, and I think, you know, this has been a big problem in education and there are a lot of professors who are very unhappy about just how much AI has invaded, you know, college campuses and it's a very, very tricky thing to To combat there's just no world in which people are gonna using it wholesale. I see this like learned helplessness, like not not just among students, but even among professionals who feel like Oh AI could write a better email than me. So like why would I ever write an email myself And I think this kind of comes from like a place of insecurity but Yeah. alsoso it's just like It's kind of concerning. Like if you have ChatTP write all your emails, then you're N going be able to write a decent email yourself Yeah. and and At the same time, I think For college students you know, if you believe that everyone else is getting A's from using ChatyPT, you know, you're in sort of this impossible arms race. But you're right, like you're not gonna to learn You only ever use these tools. And so I think having any line of defense against this stuff for professors or really any professional in a workplace who needs genuine human work is really, really, very important I'm curious whether your job gets harder as time goes on because these models change all the time and I'm curious if You think they will evolve in ways that get harder to detect Um I'm also curious, you know, you have to at Train on human writing How do you guarantee You're getting human writing and fresh human writing Yeah, Okaykay, a lot to unpack here. So I do think models have gotten much more capable in the last six to nine months. I think they've gotten much better at using their context and using a lot of context So we've gone from like a really short to chat GPT to instead people who are writing essays with clot code. L let's let's do a deep dive on the like straight of hormz and like the Iran situation. and then Claude will go, do a bunch of research, make a bunch of web calls, write down a bunch of contacts, and then compile a big essay. And so even though all of this is autonomous It's using a lot more of its context and so its output is much more detailed than you would have previously expected. And so Part of that Part of our job is u trying to understand like what AI writing looks like in this new paradigm where' where these models are much more capable And then the other side is, yeah, the human side U, I think A lot of our human data today comes from the pre twenty twenty two Internet pre chat GPT where we know for sure Um Most text was written by AI And I think there's like What we're looking at today is like, if If we're going to use a datas set from twenty twenty six, we need to be incredibly certain that it's not contaminated and there's not. AI or ChriT outputs in there. Is there a future where you are paying people to sit in a room at a notepad and handwrite essays in order to get like true human output. We've literally discussed this. We've talked about like having an essay contest or something, but we have to like supervise you I to make sure that what people are writing it are legit. I think we've landed on some like simpler, lower tech solutions, like just looking for thusted writers and sources that we believe are either not using AI or using AI and like a way which we don't think we need to detect, for example, like I'm just gonna ask Claude like like this idiom's on the tip of my tongue. And so like I'm going to ask Claude for some like wording suggestions. I think that's totally fine. and we don't want that to be flagged as AI. That's interesting. So when you say trusted sources, you mean like the New York Times or what does that look like to you Uh, yeah, I think that means like, prorofessionals who write and who have a history of writing before AI So I think like if we're looking at books, for example O. There's a lot of authors who been like been prolific and published a lot of novels before AI. And then there's some that suspiciously started putting out like three or four books a year starting in twenty twenty four and they just self publish on Amazon. And those are the ones that we want to avoid. Are you working directly with any of those authors who you trust? Not today. I think this is like really ongoing question for us of making sure U Pangram still works. as well on twenty twenty six human written content as it does on like twenty twenty human written content And so I think we're going to have some Interesting projects around that later in the year But but it's mostly just If we do our job right. Then it's invisible. thenen Then Panram continues to work and nobody really notices. So what does come next for Panram? We were chatting before we started recording and you mentioned there might be some addvancements coming soon. Yeah, so we have some interesting models that are in the oven, I think that I'm most excited for it This for our next release for our text model. which is going to do much better on humanizers. There's this whole crop of tools on the internet called humumanizers, which people can use to Basically cheat to have an AI model paraphrase your AI text in a way that it doesn't trigger an AI detector anymore. And so we've been this has been kind of an adversarial battle ourur new model is going to be much, much better here And it's also going to be better at understanding degree of AI assistance in a piece of text That's fascin It takes a remarkable dedication to laziness to make an AI essay and then run it through another AI program to hide that you ran it through AI in the first place. It's a big business actually. like there's so many of them out there and they They astrourf it and I think they're just kind of like some of the The scum of the eararth. They're the worst of the worst. So do you have to train on their outputs specifically? Yes, we do So we ran a big data collection campaign to collect a lot of data from these unionizers. And then what we did is we built our own internal humanizers. mimic what these external humanizers do So we could go from a small amount of data a large amount of data And then we train our model on it Long term, I think tools like this are going to be essential for the world to be able to tell what is human made and what is not How do you think about turning Pangram into like a real sustainable business And are you worried that one day Google goes, this is important. We're just gonna add a little button to Chrome and you know, poof, that's that You know, if Google does this, I'm happy, then you know, my work here is done and we've the problem is solved I think There's a lot of Um, competing factors here. Sort of how I think about Pangrram in general is we're building this technology, this like core infrastructure for a future where we have these pars. Powerful generative AI models If ChaGPD stopped existing and Claude and all the others then then we wouldn't have a business I think going to stick around and they're going to continue to have societal effects that we need to solve. So how do you make sure that Pengram lasts? You think that this is just going to be essential enough of a business that people will keep coming to you I think so. I genuinely think we're really well positioned because it's such a hard problem and because these
This excerpt was generated by Smart Features
All podcast names and trademarks are the property of their respective owners. Podcasts listed on Podtastic are publicly available shows distributed via RSS. Podtastic does not endorse nor is endorsed by any podcast or podcast creator listed in this directory.