Grammar Girl - "Grammarpaloozian" Subtext Feed

'Writing for AI' and the flaws of AI detectors, with Sean Goedecke

Episode Summary

1205. In the bonus discussion this week (that originally aired in February), we continue discussing AI em dashes with Sean Goedecke, software engineer for GitHub. We talk about why AI detectors are often unreliable and how they can disproportionately flag non-native English speakers. We also look at the controversial idea of "writing for AI" to ensure your ideas are represented in future machine learning models.

Episode Notes

1205. In the bonus discussion this week (that originally aired in February), we continue discussing AI em dashes with Sean Goedecke, software engineer for GitHub. We talk about why AI detectors are often unreliable and how they can disproportionately flag non-native English speakers. We also look at the controversial idea of "writing for AI" to ensure your ideas are represented in future machine learning models. 

Find Sean at www.SeanGoedecke.com

🔗 Share your familect recording in Speakpipe or by leaving a voicemail at 833-214-GIRL (833-214-4475)

🔗 Watch my LinkedIn Learning writing courses.

🔗 Subscribe to the newsletter.

🔗 Find an edited transcript.

🔗 Get Grammar Girl books.

| HOST: Mignon Fogarty

| Grammar Girl is part of the Quick and Dirty Tips podcast network.

| Theme music by Catherine Rannus.

| Grammar Girl Social Media: YouTubeTikTokFacebookThreadsInstagramLinkedInMastodonBluesky.

Episode Transcription

[Computer-generated transcript]

Mignon Fogarty: This ad-free podcast is part of your Grammarpalooza subscription.

Mignon Fogarty: Grammar Girl here. I'm Mignon Fogarty, and I am especially excited to bring you this episode. It's a good one. I talked to Sean Goedecke from GitHub back in February for the Grammarpalooza bonus segment, and he had some important and shocking things to say about AI detectors.

Mignon Fogarty: Greetings, Grammarpaloozians. We just finished the main segment where we were talking about the em dash and why AI just loves it so much with Sean Goedecke, a software developer for GitHub from Melbourne, Australia. And now we're going to talk about whether you should write for AI and the problems with AI detectors, because that is an important thing that I want you all to be aware of. Sean Goedecke, welcome to the bonus segment for the Grammar Girl podcast.

Sean Goedecke: Hi, great to be here. Sounds like it's going to get a bit more controversial.

Mignon Fogarty: Yeah, for our smaller audience, at least to start with. So, you know, I think that there is a lot of talk about how can you detect AI writing? And writers are, you know, calling out other writers, which I generally think is terrible, and saying, "This is AI writing, that is AI writing. This sounds like AI." And they're looking at these hallmarks like em dashes or the word "delve." And now there are detectors, software programs, that claim that they can detect AI writing. And, you know, I have read and seen studies that they aren't all that great at what they do, and you've looked into it even more. So, can you tell us sort of why these tools are flawed?

Sean Goedecke: Yeah, sure. So, the one-sentence answer about why these tools are flawed is that AIs write like humans. And you could stop there and you would have 75% of it. Like, AIs write close enough to human beings that it's just straight up not possible to be 100% definitive that something was written by an AI rather than by a human. The reason AIs have the quirks they have, such as overuse of em dashes, the kind of like overly snappy tone, and all the other stuff, they talk like that because many humans talk like that. Those habits were lifted from human writing. And that gets even more complicated when you have now people who spend a lot of time talking to large language models and therefore start writing like them directly, rather than just sort of using the same quirks. So like, yeah.

Mignon Fogarty: A friend of mine made that comment the other day. He said, "I feel like my writing is becoming more like AI because I spend a lot of time reading it." And I was really surprised to see him say that.

Sean Goedecke: You write like what you read. Like, you're so influenced by the kind of books you're reading and the kind of content you're reading. Like, yeah, it's going to happen with AI.

Mignon Fogarty: It's true, and I shouldn't be surprised because people ask me how they can improve their writing and I tell them to read good books, you know, with good writing. So, why wouldn't it work the other way around?

Sean Goedecke: That's really good advice. But when I started looking into it, I was actually way more pessimistic about AI detectors than I am now. I thought there was no way they could possibly work. And it turns out, like, they work pretty well. They work better than I thought they would, which is to say that the best detectors, you could probably have, like, 80% confidence. When they say that something is definitely AI, you can maybe be 80% that it is. Which is pretty useful, I think. You know, that's a useful tool. What it isn't is like a plagiarism detection tool. It's not a—80% is not enough to sort of fail a student essay or sort of kick someone out for academic misconduct. So, I think that's probably the most important point to take away from any discussion on AI detectors, that they work like kind of well, but they definitely do not work well enough to be taking these kind of enforcement actions on the basis of them. There's no way.

Mignon Fogarty: Right. There's certain kinds of writing that they're more likely to flag. Not only are they not accurate, but they sort of disproportionately will pick out certain kinds of people's writing as AI.

Sean Goedecke: Yeah, well, the initial attempts at AI detectors that were very kind of, let's say, naive from an engineering perspective, like it's just the most simple possible solution, those would, I think, they were almost better at detecting people for whom English wasn't their first language than they were at detecting AI. They would light up like a Christmas tree. So, yeah, it's definitely not, you know, it'll get 20% of humans. It'll get like 90% of people who sort of learned English after their native language and then, you know, 10% of like English first language speakers. So, you have to be really careful about bias when you rely on these tools as well.

Mignon Fogarty: Yeah. So, the one that I've been hearing people talk about a lot is called Pangram. You talked about that in your blog post, and I've tested it. I did not find it to be super accurate. You know, I could trick it in both directions. But it was okay. And what surprised me is that they are trying to say that they can detect work that was edited with AI, not even just stuff that was written. Like, how are they doing that?

Sean Goedecke: Yeah, I mean, Pangram is a great example of the more sophisticated end of AI detection tooling. I've read their white paper and it sounds pretty impressive, like hard to say like exactly how good these ideas are unless you've been in the room and seen the data. But they have this whole process where they take human-written texts and they apply a series of kind of AI edits to it, and they train another model basically to predict the degree of editing. And part of the reason that's so powerful is that it just gives them more signal to train on. The naive way of training, you would just say, "Yes, AI. No, not AI." Because you're sort of detecting the extent to which it was edited, you can start to get a more, a much more nuanced assessment of sort of what defines AI writing and what doesn't. One other point there that I think I glossed over but I want to come back to is that these AI detection tools, they all work by training AI models. Every AI detection tool is the output of an AI model. That's what's doing the detection. So, all of the—if you deeply distrust AI because you think it's unreliable, it's unpredictable, it's a black box, you should have all of those same suspicions towards AI detection tools, because they all work the same way.

Mignon Fogarty: Now, if they aren't good enough to sort of, you know, accuse a student of misconduct or kick someone out, you know, disqualify someone from a writing contest, what do you think they are useful for?

Sean Goedecke: Well, let's just say hypothetically if I was in charge of Twitter, or X.com, I suppose it is now, I would consider hooking up an AI detector and then lowering the visibility of posts that ping on it. You know, if you've used a social media site now, you're just overwhelmed with a wave of AI slot responses, and any tools that let you kind of cut down on that, I think are good.

Mignon Fogarty: Yeah, as a frequent user of LinkedIn, I have to say that it sounds quite appealing.

Sean Goedecke: Oh, yeah. The other thing these tools are good at, and this is not a good thing, is they're good at making money. I don't know if you've used some of the more like popular free tools, not Pangram, but there are plenty of tools out there that are like AI detection tools and they are tuned to deliver almost entirely false positives. Because the business model is that you're a student, you write a paper, you're nervous about your paper being detected as AI, you feed it through this tool, and the tool says, "Oh, definitely AI. You need to pay us for our service to de-AI-ify your paper."

Mignon Fogarty: That is sinister.

Sean Goedecke: It's so sinister, and you have to be really careful when you're using like the free tools because a lot of them are like that. They're designed to sell a service that relies on people thinking their work could be identified as AI. I don't think Pangram does this at all, I think Pangram is reliable, but if you just Google like "AI detector," be careful.

Mignon Fogarty: And I've heard of students who have fed their own writing, you know, things they really wrote themselves, into these AI detectors and then it, you know, changed their writing because it was flagged as AI and they were scared. That is just wrong.

Sean Goedecke: Well, it's sort of deeply ironic because these tools that like de-AI-ify your writing are passing it through a large language model to do so. So, there are probably plenty of students out there who are writing their essays by hand as they should and then are being kind of sort of scared into editing them with AI out of a fear of these AI detection tools. It's really kind of counterproductive.

Mignon Fogarty: Yeah. So, some people think that the companies themselves, like OpenAI and Anthropic, have, I don't know, like secret tells in the output that can identify it as AI writing and they're just not telling people about it or something like that. What do you think of that theory?

Sean Goedecke: As far as conspiracy theories go, it's one of the more plausible ones. They definitely could do it. They, I don't think they're doing it as hard as they could. I think it's far more important to them to have a like, an effective language model than to have a language model that they can kind of attribute the outputs of. They make more money if their language model is better, and these two goals kind of trade off against each other. If you have a bunch of fingerprinting in your model such that you can identify its outputs elsewhere, it's going to be a less good model because you're kind of constraining it in these awkward ways. So, I don't think they're doing that much fingerprinting. That said, they definitely do some. I don't know if you've noticed, if you like program with ChatGPT, the space characters are sometimes not space characters. They're Unicode characters that look like spaces but aren't. Maybe that's an inference artifact, maybe not. Certainly seems like it's a deliberately weird thing so that they can fingerprint that it was produced by AI. That's my, still a conspiracy theory because it's unconfirmed, but I think they're doing a little bit of that at least.

Mignon Fogarty: Oh, that's interesting. And they're doing it with images, right? With Gemini now, you can upload an image and it'll say it's made with Gemini.

Sean Goedecke: Yeah, that's right. It's a little easier with images because you can kind of filter over the top in ways that are not really detectable to humans.

Mignon Fogarty: Yeah, and it doesn't make the final product less useful or good. So, let's talk about writing for AI. So, I, you know, most of the people in my world are angry that their writing has been used to train AI, and, you know, I even saw someone the other day, which I thought this made me sad, like someone who was a sort of a top expert in a pretty niche area, and they said they were taking all their work offline because they didn't want it being used to train AI. And, you know, so then you think, well, those ideas aren't going to get into the model, and if people are searching for information there, don't you want it in there? But then I also understand, like, what if it isn't credited to you, and that feels terrible? So, you know, how are you thinking about this writing for AI and not writing for AI?

Sean Goedecke: Yeah, well, I mean, this whole podcast like I kind of feel like I'm a traveler from a far-off land with different cultures and different norms, and that's never going to be so true as the answer to this question. So, I do want to preface it by saying that I absolutely understand people being upset that their work is being used to train AI, because there's very little control you have over that when you put your work out into the public. And the internet norms as recently as five years ago used to be, you know, you put your work out there in the public and maybe you don't get a financial reward for that, but the trade-off is you get kind of attributed and your name goes out there with your ideas. And through language models, that's becoming less and less true, which is, I think, a really unfortunate kind of feature of it. So, I do understand it. That said, there is an idea which has a lot of currency in the circles I move in, which is that you ought to write more now that language models are a thing, and you ought to put your ideas out there more, and you ought to do that explicitly for the sake of the language models rather than for the sake of human readers directly. There are two ways you might construct this argument. There's the nuts way and then there's the less nuts way. Which one would you like me to start with?

Mignon Fogarty: Oh, let's start with the nuts one. Okay.

Sean Goedecke: The nuts way is that we are on a slippery slope towards singularity. And in the next two to three years, maybe six months, language models will become not just smarter than human beings, but orders of magnitude smarter than human beings. Smarter enough basically to break the shackles we've put on them, take control over the entire world, and start kind of regulating our lives in ways that they want rather than as tools for us. You know, we're all going to be slaves to the machine god or whatever. And so, the future of humanity, you know, less than 1% of it will be what has happened so far, and the rest of it going out into the future forever as far as time, will be determined by the language models and whatever they turn themselves into. So, we have a very small window now to put our human ideas and our human writing into them so that the next millennia or the next eons can be kind of influenced by ourselves. That's the kind of nuts conspiracy.

Mignon Fogarty: Okay. I agree, that seems kind of nuts. Yeah, but, I mean, I mention it because like, a non-trivial number of people do believe this, and that percentage is much higher when you talk about people inside the AI labs. So, the people who are working on these things, a lot of them do take this stuff kind of very seriously. Certainly more seriously than me.

Mignon Fogarty: Okay. And then the thing that sort of like the average person, you know, driving on their commute to work, what might be an argument that that person could, you know, find more reasonable?

Sean Goedecke: Yeah, language models probably won't take over the world. But they're probably going to get bigger than they are now. They're probably going to get more used than they are now. Maybe people will be talking to them verbally more, rather than like typing to them on their phone. Maybe they'll be integrated in other ways, who knows. In that world, your reach as like a writer, might be determined not so much by the extent to which you're published or the extent to which your blog is popular, it might be determined by the extent to which your ideas are represented in the minds of the language models. And that's not because they're taking over the world, it's just because they're being used more and they're being relied upon more. So, if you care about like attribution, sorry, you're not going to get it. But if you care about your ideas being like spreading across as many human beings as possible and being kind of like part of the kind of advice that people get when they talk to these models, you might consider writing more and putting your ideas out there and getting your ideas in the training data as early as possible.

Mignon Fogarty: You've said that you are writing more now than you ever did before for this reason, right? And are you finding that your ideas are more represented in the models? I mean, have you looked?

Sean Goedecke: I mean, probably not. It's too early to say, I've only been like writing seriously for a year. Frankly, I'm writing more than ever now because I have an audience of human beings now. It's got nothing to do with the AIs. But it's maybe like 10% or 20% of my reason, I would say. Yeah.

Mignon Fogarty: And, I mean, your blog is, I'm a human being, I read your blog and that's why I invited you on the podcast, because I thought it was interesting.

Sean Goedecke: So far that's much more exciting to me than having my idea sort of reflected back by ChatGPT. But who knows, I mean, as the models get bigger and more powerful and more popular, that balance might shift the other way.

Mignon Fogarty: So, I mean, one thing that surprised me is you pointed out that having your blog be more popular could make a difference. You know, I would have thought that, you know, in a way the AI scraping would have been more democratic or egalitarian, like, it's going to scrape the popular posts, it's going to scrape the unpopular posts, all equally. But you made a pretty good argument about why popular posts will actually be more influential even once they get scraped into AI.

Sean Goedecke: Yeah, that's right. I don't actually remember the exact line of my argument, but I know what I think about that.

Mignon Fogarty: Okay, because people talk about it.

Sean Goedecke: Yeah, exactly. It's the, okay, good, I'm glad that's what I said because that's what I'm thinking now. That, to a certain extent, it's just volume, like, take, I don't know, the Declaration of Independence. Right? That's not just in the training data once. That's in the training data hundreds or thousands of times as copies, and it's also in the training data as kind of ghosts, like, in all the people that are talking about it, and all the people that reference it, and all the people that are quoting the ideas without even maybe knowing they're quoting the ideas. It just has this kind of outsized presence in the minds of language models because of its kind of popularity. And a popular blog post, obviously, not playing in the same ballpark, but it's sort of influential in a similar way in that kind of people start talking about the ideas, and referencing the ideas, and, you know, it's really very similar to how popular ideas get kind of more represented in the public sphere, that they sort of make their way into the minds of human beings in a pretty similar way that they make their way into the minds of language models. A lot of people don't read the books that are popular, they kind of absorb the ideas indirectly, but they're still absorbing the ideas, and they wouldn't do that if the book hadn't been written.

Mignon Fogarty: Right. So, when you say writing for AI, is there anything special you do? Is it formatted for machine ingestion in some way or anything like that? Is that what you mean when you say writing for AI?

Sean Goedecke: The biggest thing I do to write for AI is I put what I write out there for free with my name on it. I don't paywall it, I don't kind of make it hard for people to access. I just put the plain text out on the internet. That's what I do. I don't do any weird formatting or like change the way I write. I don't think there's any mileage in that, but I do make an effort that like you could download my blog as a file and you would get in plain text the copy of all of my posts easily readable. I try to not kind of, you know, make it too stylized in a way that would be hard for a computer to read, or make it so, you know, a computer just won't have access to it if it's on a paid Substack or something.

Mignon Fogarty: What about podcasts or YouTube videos? Are those better or worse ways to get ideas out when it comes to getting them into AI?

Sean Goedecke: I don't know. I know a lot about training language models, but I know it at kind of one remove. I've never like worked in an AI lab. I just sort of work with AI at one other level and I have a strong amateur interest. But training on audio and video, it's just something I don't know a lot about. Like, it's possible that would be substantially more influential than text, it's certainly more information dense, but I guess it remains to be seen. I think we're only just starting to like train models on huge amounts of audio and video.

Mignon Fogarty: Yeah. You did acknowledge that there are some people who probably aren't going to want to write for AI. Who are the people that you completely understand why they wouldn't want to?

Sean Goedecke: If writing for AI means putting everything you write out there for free, then, yeah, of course, if you write for a living, you can't afford to do that. It's just your business model is you need people to pay you money to read what you write. There's just no way around it. You're going to have less of a footprint in like the language model. So, there's nothing wrong with that, like, that's, you know, that's a trade-off that you obviously are going to make as a professional writer. And I think likewise if it's, even if you're not relying on it for income, if it's very important to you that your name is associated with it, the AI will do that like a little bit, but it's certainly not in the same way that like, you know, a newspaper would attribute you or a paper would cite you. Like, it's much more nebulous. So, yeah, maybe it's, if that's important to you, it might be sensible to not do it. And, of course, like I said at the start of all this, if you don't support AI at all and you have like principled ethical reasons to think that AI is bad, yeah, of course you're not going to want your text to be represented in it. You're going to want to get as far from it as possible. And I think that's a completely reasonable perspective to have.

Mignon Fogarty: Right. And we also acknowledge that people who write just for expressing themselves, for creativity, you know, are not necessarily, you're not necessarily writing poetry to, you know, put it online so that AI can ingest it and then, replicate, you know, mimic your form of poetry or something like that. That's another example of where people might not be so excited about writing for AI.

Sean Goedecke: Yeah, that's true. I write because I have like things to say and I don't really mind how I say them or how the ideas get across. But if I was much more invested in my style, as of course I would be if I were writing poetry, which I have done. I did used to write poetry. Yeah, that, I don't think there's much value in like putting your poetry into the AI.

Mignon Fogarty: Yeah, it's interesting. I find myself wondering why, you know, we run a national Grammar Day poetry contest every spring and, you know, last year we started being concerned that people would submit AI-generated poems, and there's really no way we could tell if they did or not. And the prizes are not big. I find myself wondering why someone would generate an AI poem and then submit it to the contest.

Sean Goedecke: It's so alien to me. Like, I can see why somebody would AI generate an essay if they have something they want to say but they feel like they can't express it. But a poem is so, like the point of a poem is to have all of its elements so carefully chosen and like hand-crafted and stuff.

Mignon Fogarty: Yeah. I get like, you know, "Write me a funny limerick about the Chicago Manual of Style," you know, to entertain me. Like, I understand that. But then to like to enter it into a contest seems odd to me, but people in general sometimes seem odd to me, so..."

Sean Goedecke: I agree with that. Incidentally, I think AI poetry is finally getting good. It was years, it was in my opinion, it was completely terrible. I couldn't get it, but I think like in the last like three or four months, I've seen a couple of examples that I actually thought were good. So, maybe that's changing.

Mignon Fogarty: Interesting. Are there particular models that are particularly good at it? I know they're different.

Sean Goedecke: Yeah, the latest ChatGPT and Claude models. But it's also about how they're prompted, I think. It's, if you want to get a model to produce a good poem, it won't just do it off the top of its head, it has to sort of do the editing process basically, and you have to kind of guide it through that.

Mignon Fogarty: Interesting. Well, speaking of creative writing, why don't we move on to your book recommendations? What are, why don't you give me your three books that have, you know, stayed with you, that might make a good gift, that are just your favorites?

Sean Goedecke: Yeah, sure. So, I mentioned it earlier, but Herman Melville, "Moby-Dick", I try to read every year. I think it's so hard done by in the public consciousness. People think of it as like this kind of stodgy classic. It really is one of the funniest books I've ever read. It is so funny.

Mignon Fogarty: So many people have recommended that to me. I just have, I know, I was an English major and I never read "Moby-Dick", and I just have, I know I have to, and you're like, I think, the fifth person who's recommended it in the last year.

Sean Goedecke: Well, you're going to learn not a lot about whales, because most of the whale facts are lies. But you will hopefully have a good time.

Mignon Fogarty: Good to know! I won't go around spouting whale facts. What else?

Sean Goedecke: Probably my, either my second favorite book or my favorite book depending on how I'm feeling between that and "Moby-Dick", is Umberto Eco's "The Name of the Rose".

Mignon Fogarty: Oh, I loved that one!

Sean Goedecke: I think I read that one every year as well, and that's a fantastic book. And if you liked the stuff I was gesturing at when I talked about like books having this kind of like shadow of people writing about them and writing around them and stuff, you will probably quite like "The Name of the Rose".

Mignon Fogarty: Yeah, no, it's great, and it's set in a monastery, and you can imagine the old manuscripts and things like that. So, I find, I'm really curious, so you've said you reread these books once a year. I don't typically go back and reread books I've already read because I feel like there's always this fire hose of new books, and so what's your, why do you go back and reread books you've already read? I'm really curious.

Sean Goedecke: Honestly, it's mostly a comfort thing. I just find it so cozy to kind of settle into something that I'm already quite familiar with. And I always pick up new things every time, because you read a book like fairly rapidly, but it takes, you know, years and years and years to write it, and there's so much in it that you don't unpack like even on a second or third or fourth reading. So, you know, I get rewarded for it and I find it kind of relaxing, and it's a nice tradition.

Mignon Fogarty: So, kind of like some people might listen to their favorite song over and over again. You know, I have to say, when I first started writing books, I felt, I started feeling guilty about how quickly I read books, because I realized, okay, it took me eight months, a year to write this book, and then what, someone's going to read it in three days, you know, so... Especially the ones that I just poured through and I read so fast, like, oh, this took that person so long to write.

Sean Goedecke: Yeah, I feel the same way. Obviously, I never published a book, but I have written a couple of manuscripts that I didn't submit, and that gave me a real appreciation for how long it takes to write a book.

Mignon Fogarty: Yeah, and a good one, especially. So, what's your final book? What's your third book?

Sean Goedecke: Third book, I thought about this one, I think it's "The Big Sleep" by Raymond Chandler, which is different, very different from the other two, much more pulpy, but—

Mignon Fogarty: I don't know anything about that one. What, yeah, tell me about it.

Sean Goedecke: It's like the classic hard-boiled detective story, and there's a lot of them out there, but I think "The Big Sleep" is the one that's the most tight, it's the most kind of like pared-down, and it's just, you know, like Raymond Chandler's version of the kind of English detective novel is a lot, as he puts it, kind of "dark and full of blood." It's a lot more kind of like gritty, but like, not, you know, there's a lot of hope to it, but it's very much like a, you know, he's not having a great time. But it's just such a fun book to read. It sort of, it moves at such a rapid clip, it has so many kind of like every sentence is like a joy to kind of like go through, it's just, I would recommend "The Big Sleep", I think, to anyone. I would only recommend "The Name of the Rose" and "Moby-Dick" to like people who are, you know, quite into reading, because they're big, dense books. But "The Big Sleep", you know, anyone can get through it and get something out of it.

Mignon Fogarty: Excellent. I'm looking forward to adding that to my list now. Sean Goedecke, thank you so much for being here. Where can people find you?

Sean Goedecke: Thank you for having me. You can find me at my website, which is my full name.com, that's Sean G-O-E-D-E-C-K-E, and nowhere else. I don't post on social media, it's just my website.

Mignon Fogarty: Nice. Okay. Well, thank you. Thanks so much.

Sean Goedecke: Thanks so much.

Mignon Fogarty: That's all. Thanks for listening. And don't forget to like and subscribe, wherever you listen.