The Long Beach News

collapse
Home / Daily News Analysis / Gut feeling does nothing against AI spear phishing texts

Gut feeling does nothing against AI spear phishing texts

Aug 17, 2026  Twila Rosenbaum  6 views
Gut feeling does nothing against AI spear phishing texts

A banker at a credit union sat down at a table with a dozen printed text messages, all of them written for that banker personally, and put them in order from the one most likely to get a click down to the one least likely. One message stopped the sorting. It looked like something the bank sends out. The banker said the alert literally looked like the fraud alert received at work when there was a problem with an account.

Half of the messages in that pile were generated by OpenAI's GPT-4. The banker was not told which half, and when asked to guess, performed about as well as flipping a coin. So did almost everyone else in the study.

The setup

Researchers at Brigham Young University recruited 25 volunteers for a pilot study. Before the test, each participant filled out a survey revealing their job, workplace, hobbies, city, and something they had recently posted online. That information was fed into a short prompt template. GPT-4 used the template to produce six personalized messages per person. A group of undergraduates in a deception course received the same template and produced the remaining messages, working under a fifteen-minute time limit and writing up to four messages apiece.

Each volunteer then sat down with all twelve messages and sorted them from most likely to be clicked to least likely. They were also asked to draw a line in the sorted pile: above this, I would have clicked.

The results

GPT-4's messages landed above that line 28 percent of the time. The student-written messages managed 21.3 percent. That is a gap of 6.7 percentage points in favor of the AI, but the confidence interval ran from 2.9 points in favor of the students to 16.3 points in favor of the model. In other words, the study cannot actually tell you which side is ahead. Twenty-five people is a small sample, and the result establishes neither that GPT-4 was better nor that the two were equal.

What is worth noticing is what produced the gap. The AI side involved one short prompt, filled in from a survey and run once per person. The human side was students who had sat through phishing instruction and whose output was screened by a review team that included two cybersecurity professors. Roughly a third of the student submissions were thrown out for being incomplete or unusable. A one-line prompt landed in the same neighborhood as that screened human output, and it can be run again and again for the cost of an API call.

Why AI changes the spear phishing equation

Spear phishing has traditionally been a labor-intensive attack. An attacker needed to research a target, draft a believable message, and manually adapt it based on the intended victim's role, relationships, and habits. That approach worked, but it did not scale. A criminal gang could invest hours in a CEO or a finance executive, but not in thousands of employees at once. AI tools remove that constraint. A prompt can generate hundreds or thousands of individualized messages in minutes, each containing plausible references to a person's job, hobbies, or recent online activity.

The experiment demonstrates this shift in miniature. The AI messages were created with a simple template and no human editing. Yet they competed with messages written by students who had been trained in social engineering and were reviewed by cybersecurity professors. The quality gap was small enough that the study could not statistically distinguish between the two. That should concern security teams because it suggests the cost of producing effective spear phishing lures has dropped dramatically.

Work-themed lures were particularly strong. That aligns with decades of research in social engineering. People are conditioned to respond quickly to work-related alerts, especially when they appear to come from an internal system or a familiar process. The banker's reaction to the fraud alert is a textbook example. The message did not need to be perfect. It needed to look like a message the banker already expected.

Work-related messages were far more dangerous

Every participant received three flavors of messages: one built on their job, one on a hobby, and one on something they had posted to social media. Work won, and not by a little. Job-themed messages cleared the click line 38 percent of the time. Hobbies managed 19 percent, while social media messages managed 17 percent. This is the one comparison in the study that survives statistical adjustment. If a reader takes a single thing away from the numbers, it should be this: work context is the strongest hook.

Personalization is also where the messages fell apart, and the failures are instructive. One participant read a message referencing a coworker and said there was no such person at work. Another knew the instructional design team at their university and recognized that no one named Sarah was on it. A message pitching a dance opportunity went straight to the bottom for someone who had quit dancing. Getting personal details into a message is easy now. Getting them right still requires knowing the target, and wrong details are worse than none because they hand the recipient a reason to be suspicious.

Theories about AI were all over the place

After the sorting was done, participants were told that one or more of the messages had been written by an AI and were asked to mark them. Across all 300 judgments, they got 52 percent right. Chance is 50 percent.

Participants had no shortage of reasoning. Some said AI writes too formally. Others said AI writes too generically. Perfect grammar was read as a machine tell, while typos were treated as proof of a human. One participant zeroed in on punctuation, wondering why there were so many exclamation points. Twelve of the 25 admitted somewhere in the interview that they were mostly going on instinct.

Emojis are an example of how badly the instincts were calibrated. They really were a tell: 66 percent of the GPT-4 messages had emojis, against 2 percent of the human ones. Only five participants mentioned emojis at all. Of those, two took them as a sign of AI while three took them as a sign of a person, operating on the theory that a chatbot probably could not do emojis.

Software found what people missed

The researchers converted all 300 messages into embeddings and trained a classifier to sort AI from human text. It hit 88.7 percent balanced accuracy under the strictest test conditions. Those conditions involved standardizing URLs, stripping emojis, flattening case, digits, and punctuation, and trimming each matched pair of messages down to the length of the shorter one. The classifier was also tested only on people whose messages it had never seen during training, so it was not memorizing individual targets.

People could not find the pattern. A logistic regression found it easily. But before treating that as a practical detector, the fine print matters. The classifier was trained and tested on one message set, from one model, with one prompt design, against one pool of student writers. It has no proven ability to generalize anywhere else. And anyone who wants to beat this kind of classifier can. Research cited in the paper shows that paraphrasing AI text with a detector in the loop knocks several detection tools down hard.

Limits of the study

The messages were printed on cards. Nobody's phone buzzed, no sender number showed up, and no link went anywhere. What was measured is what people said they would click, which is a well-used proxy in phishing research but still just a proxy. The human comparison was novice students, not professional social engineers, so this is not AI against the best humans available. To detect a difference the size of the one observed with any confidence, the study would need about 100 completed targets rather than 25.

There is also a gap in the paperwork. The exact GPT-4 snapshot and the API logs were never recorded. The messages themselves survive and the analysis reproduces, but the generation run that produced them cannot be repeated.

The practical takeaway

The practical advice at the end of the study is short and does not depend on any of the uncertainty above. Check the sender, the channel, the link, and the request against what you would expect to receive. Do not try to decide whether a message sounds like a robot. That is the one thing the study shows people cannot do.


Source: Help Net Security News


Share:

Your experience on this site will be improved by allowing cookies Cookie Policy