Two JAFF-reading acquaintances of mine, one broadly pro-AI and the other fairly AI-neutral, have both complained recently about the AI-written JAFF turning up in Kindle Unlimited. The ideas are often interesting. The prose is so tedious and repetitive that many of these are barely readable and absolutely unlistenable. Audio is a merciless format: it exposes every clank and clunk in a sentence, along with any typos.
For the love of Heaven, AI authors, please use stylometry to get your skynets writing with a little more human rhythm. Stylometry is the statistical analysis of literary style: sentence length and how much it varies, sentence structure, vocabulary, punctuation habits, and so on. Here’s how to get started.
- Install spaCy (spacy.io) in your Python environment, plus a language model such as
en_core_web_sm. spaCy is a general-purpose language-processing library, and it handles the sentence splitting and part-of-speech tagging you need. - Ask your AI to write a Python script that reports the statistics you care about: mean and spread of sentence length, commas and semicolons per thousand words, dialogue versus narration, and so on.
- Download a batch of public domain texts from Gutenberg by an author you like. Strip the Gutenberg license text from the front and back so it doesn’t muddy the numbers, then run the script on the lot.
Georgette Heyer’s Regency novels are not public domain, but Gutenberg has four of her early novels, and that’s enough to get started. They’re historicals from various eras rather than her mature Regency style, so they won’t help much if you’re trying to imitate her voice. They’re plenty for the statistics, though. When I ran a stylometric analysis on my own work, the numbers came out surprisingly consistent across first- and third-person books and every genre and setting I’d written in.
Even on a bad day, Heyer was probably a better prose stylist than you or me, and certainly a better one than our clanker friends. There’s a bonus, too. Gutenberg’s txt files render em-dashes as double hyphens, which your script won’t count as em-dashes, so the benchmarks come out em-dash free. That helps offset one of the AI’s most notorious tics.
Once you have the analysis, take the numbers back to your AI and explain that these are rough benchmarks for its writing style. It doesn’t need to hit them in the initial composition. It should audit its drafts against them afterward and rewrite where it misses. Stylometry won’t cure everything, since stock phrases and repetitive vocabulary need their own attention, but it goes a long way toward fixing the rhythm.
For a more detailed walkthrough, please see this Future Fiction Academy video on the method:
Then run the revised chapter through text-to-speech. If it still clanks, try, try again.
