A long time ago, when walking around Amsterdam with a gay friend during Gay Pride, the subject of the Gay Radar (gaydar) came up. Are some people better than others in ‘detecting’ gay people based on their appearance and behavior. My friend said something so simple that forever changed my perspective on the world: “You only see the ones that are obvious to you.”1
Let’s say that out of 1,000 random people you meet, 30 are gay. Let’s further say that you thought 12 were gay based on what you think a gay person is like. You ask all 12 (this is thought experiment, ok 😂) if they are really gay, and 11 confirm. Wow! you got 11 out of 12 right, that is 92% accuracy. What an impressive gaydar. But what about the 18 you missed? 11 ‘detections’ out of the 30 total gay people you met means only 37% ‘detection’. A lot less impressive.
In data science, the 92% is called precision: From the gay people you detected, what % of the time were you right. The 37% is called recall: what % of the true total number of gay people did you detect. It is very, very, very easy to fool ourselves into thinking we are very good at detecting something if we know we have high precision but don’t know the recall. And, it is actually not easy to know the recall level. It requires backward looking sampling and usually we are not in a position to that.
Detecting AI Generated writing
It is fun to call out AI Generated writing and be correct. I am very good at it, and so are many others I know. The reason I’m good at it is that I’ve sent something like 20,000 prompts to these models through the course of making Magicdoor, for work, or just for fun. If someone does not take any steps to change the default writing style of any of the mainstream AI models, I will spot the output from a mile away. I am almost never wrong. My precision is very high.
But I can also prompt AI models to output text that will score a full-on 0% ‘AI Suspicion’ on ZeroGPT. A while back I plugged all of my Substack writing (all of which is hand-written) into AI to repurpose it into LinkedIn posts. Every one of those posts (one, two, three) passes AI detectors as human written.
Then there is of course Reddit, where people are sharing prompt tips and tools to get AI text to escape detection. This is not hard to do and therefore I am sure my recall is very low. So if you think you are a ChatGPT sniper and that you can reliably sniff out when it’s being used, think again.
AI Detectors are pointless
There are tons of reports online of students getting failing grades because their teachers used something like ZeroGPT and detected AI. On the other hand, as I showed above, ZeroGPT is easily defeated with just a little bit of effort. It’s also possible, although I find it much harder, to do it the other way around and hand-write text that is flagged as 100% likely to be AI. This is a snippet that I hand-wrote that got a 100% AI likelihood score on ZeroGPT:
Because this was quite hard to do, I think (but cannot be sure) ZeroGPT (again) has a a fairly high precision. So with medium confidence I am willing to say that not very many students are unfairly getting failed for their human written essays being flagged as AI. Are there going to be some people who’s natural writing looks very AI generated? Statistics suggests there have to be. But probably not very many. On the other hand, I am much more confident in saying that AI Detectors are completely useless as a way to catch a high share of AI usage (low recall!). Everyone who knows their submission will be checked with an AI Detector will be able to prompt their way around it with very low effort.
What then?
For middle and high school, two things:
Students will need to learn to get good at using AI anyway. Let them use AI however they want for essays written at home. Teach them how.
Grade the essays on quality of arguments, clarity, originality — things AI tends to be very bad at.
Task students with using AI to write a play in the style of Hamlet but about current events (for example).
To learn to think and write well, have students write in school, on air gapped devices or with pen and paper.
A product idea for HumanTyped Certificates
More generally, I’ve been walking around with the idea of Certified Human Typed content for a while:
Create a text editor, which has no way whatsoever to paste content. The only way to get content into the document is to type it. If you really wanted to use AI, you’d have to type it over letter by letter.
For content created with this editor, provide a certificate that it is HumanTyped, potentially with a crypto element (immutably stored as a record on a blockchain).
Customers could be schools. Probably only schools. Unless people start challenging authors more and more on proving that their writing is not AI generated. But I am not sure that people would do that… Not much more than food for thought at this point. But anyway, here is a quick prototype: Try it out!
I could come up with a synthetic made up example to avoid talking about detecting gay people, but this is genuinely what happened to me personally. This was the AHA moment for me to intuitively get the idea of recall, without even knowing the word for it at the time. For the record, I don’t believe it is important or in any way necessary to be able to ‘detect’ gay people. Everyone should be free to love whoever they want.



