BIP NYC

collapse
Home / Daily News Analysis / Gut feeling does nothing against AI spear phishing texts

Gut feeling does nothing against AI spear phishing texts

Aug 17, 2026  Twila Rosenbaum  12 views
Gut feeling does nothing against AI spear phishing texts

A banker at a credit union sat down with a dozen printed text messages, each one written specifically for that banker, and tried to put them in order from most likely to get a click to least likely. The sorting stopped at one message. It looked exactly like the fraud alert the bank sends internally. That message, it turned out, was generated by an AI.

The moment comes from a pilot study at Brigham Young University designed to test whether people can trust their instincts when facing AI spear phishing messages. The answer, according to the researchers, is no—at least not in any way that matters.

25 volunteers, 12 text messages, one uncomfortable question

The study brought in 25 volunteers. Before the test, each one filled out a survey that gave up details about their job, workplace, hobbies, city, and something they had recently posted online. Those details were fed into a short prompt template. GPT-4 used that template to generate six personalized messages per person. A pool of undergraduates in a deception course worked with the same template to produce the rest, under a fifteen-minute clock for up to four messages apiece. The student-written messages were then screened by a review team that included two cybersecurity professors. Roughly a third of the submissions were thrown out for being incomplete or unusable.

Each volunteer then sat down with all twelve printed messages and sorted them into a pile. They drew a line in the pile: above this, I would have clicked. It was a simple setup, but it produced a sobering result.

AI-generated messages held their own

GPT-4's messages landed above the click line 28% of the time. The student-written ones managed 21.3%. That gap of 6.7 percentage points looks like a win for the AI model, but the confidence interval tells a more careful story. The interval runs from 2.9 points in favor of the students to 16.3 points in favor of the model. In plain terms, the study cannot establish whether GPT-4 was actually better. Twenty-five people is a small sample, and the overlap in uncertainty means the result is inconclusive on that specific question.

What matters more is how the gap was produced. The AI side needed just one short prompt, filled in from a survey, run once per person. The human side required students who had sat through phishing instruction and a review team that discarded about a third of their work. A single prompt fell into the same statistical neighborhood as that heavily curated human effort—and it can be run again, and again, for the cost of an API call. That asymmetry is where the real danger lives.

Work-related lures were far more dangerous

The researchers tested three flavors of personalized message: one built on the participant's job, one built on a hobby, and one built on something they had posted to social media. Work won, and it was not a close contest.

Job-themed messages cleared the click line 38% of the time. Hobbies managed just 19%, and social media posts 17%. This was the one comparison in the study that survived statistical adjustment. If a single number has to be taken from this research, it is this one: when a spear phishing message feels work-related, people are much more likely to click.

That makes sense from an attacker's perspective. Work-related messages arrive with built-in context. They can mimic internal alerts, vendor notifications, or messages from colleagues. They also create a sense of urgency that personal lures often lack. A bank warning about fraud, a message from IT, or a request from a manager all press the same psychological buttons: fear, duty, and obligation.

Personalization cuts both ways

Personalization is also where the messages fell apart, and the failures are instructive. One participant read a message that referenced a coworker and said, “there's no Mike at work.” Another knew the instructional design team at their university and knew there was no Sarah on it. A message pitching a dance opportunity went straight to the bottom of the pile for someone who had quit dancing.

These moments show both how strong and how brittle AI-generated personalization is. Getting personal details into a message is easy now. Getting them right still requires knowing the target, and wrong details are worse than none because they hand the recipient a concrete reason to be suspicious. The technology is improving quickly, though. The gap between a convincing reference and an embarrassing mistake will narrow as AI tools gain access to better data and more context.

Tech instincts were wrong almost every time

After the sorting was done, participants were told that one or more of the messages had been written by an AI and were asked to mark them. Across all 300 judgments, they got 52% right. Chance is 50%. In other words, they might as well have flipped a coin.

That did not stop them from having theories. Some said AI writes too formally. Others said AI writes too generically. Perfect grammar was read as a machine tell, while typos were considered proof of a human. One participant zeroed in on punctuation: “why are there so many exclamation points all over?” Twelve of the 25 admitted somewhere in the interview that they were mostly going on instinct.

The emoji example shows how badly those instincts were calibrated. Emojis really were a tell: 66% of the GPT-4 messages had them, against only 2% of the human ones. Yet only five participants mentioned emojis at all. Two took them as a sign of AI, while three took them as a sign of a person, on the theory that a chatbot probably could not do emojis. The actual pattern was almost the opposite of what intuition suggested.

Software found what people missed

The researchers did not stop at human judgment. They converted all 300 messages into embeddings and trained a classifier to sort AI-written text from human-written text. It hit 88.7% balanced accuracy under the strictest test conditions available.

Those conditions mattered. The text was standardized in several ways: URLs were normalized, emojis were stripped, case, digits, and punctuation were flattened, and each matched pair of messages was trimmed down to the length of the shorter one. The classifier was also tested only on people whose messages it had never seen during training, so it was not memorizing individual targets.

People could not find the pattern. A logistic regression found it easily. That is an important finding for security teams, but it comes with fine print. The classifier was trained and tested on one message set, from one model, with one prompt design, against one pool of student writers. It has no proven ability to generalize anywhere else. And anyone who wants to beat this kind of classifier can do it: research cited in the paper shows that paraphrasing AI text with a detector in the loop knocks several of these tools down hard.

Limits of a Tuesday-afternoon experiment

There are several reasons to be careful before treating this study as the final word. The messages were printed on cards. Nobody's phone buzzed, no sender number showed up, and no link went anywhere. What was measured is what people said they would click, which is a well-used proxy in phishing research but still just a proxy.

The human comparison was novice students, not professional social engineers. This is not AI against the best humans available. To detect a difference the size of the one observed with any confidence, the study would need about 100 completed targets rather than 25. There is also a gap in the paperwork: the exact GPT-4 snapshot and the API logs were never recorded, so while the messages themselves survive and the analysis reproduces, the generation run that produced them cannot be repeated.

What this means for everyday defense

The practical advice at the end of the study is short, and it does not depend on any of the uncertainty above. Check the sender, the channel, the link, and the request against what you would expect to receive. Do not try to decide whether a message sounds like a robot. That is the one thing the study shows people cannot do.

Work-related messages deserve extra scrutiny. If an alert references your company, your team, or a familiar internal process, verify it through a separate channel before clicking. The same caution applies to messages that mention a coworker by name, a project you are known to be working on, or a recent post you made online. The data AI can use to personalize a phishing message is growing every day, and the cost of generating a thousand variations is nearly zero.

The broader lesson is that human intuition is not a useful line of defense against AI-generated phishing. Security awareness training that tells people to trust their gut is setting them up for failure. Instead, organizations should focus on structural controls: multi-factor authentication, email and SMS filtering, URL inspection, and clear reporting procedures. Those controls do not depend on a person's ability to hear the hidden sound of a language model in a text message.

People will keep looking for tells: a strange phrase, an overused word, a punctually perfect sentence. The study suggests those tells are unreliable or already gone. By the time someone can point to a reason a message feels artificial, the message may already be in production across thousands of inboxes, each one tailored to a different target. The only reliable defense is a process, not a feeling.


Source: Help Net Security News


Share:

Your experience on this site will be improved by allowing cookies Cookie Policy