https://www.smithsonianmag.com/arts-culture/review-of-the-pr...
Where do you think LLM's learned these things from? They are widely used in literary writing. Like magazines and books.
If anything, the length of that article shows how rarely em-dashes were used by most writers. They're like exclamatory versions of semicolons, a contrived sudden interruption, a sort of inversion of the three dot "…" elipsis. Maybe the em-dash cracked and fell on the floor.
The reason LLMs use a lot of em-dashes is because that's a format they've chosen for output. Thinking that LLMs have a lot of em-dashes because works in the wild have a lot of em-dashes is like thinking that LLM output has a lot of emoticons because a lot of essayists use emoticons to mark subject divisions in the text.
There are also em-dashes in a huge number of their articles. I didn't spend time picking one. I just went back to the oldest article in the first category I picked, and found one on the first try. It's a common style for more "serious" magazines and always has been.
> Thinking that LLMs have a lot of em-dashes because works in the wild have a lot of em-dashes is like thinking that LLM output has a lot of emoticons because a lot of essayists use emoticons to mark subject divisions in the text.
No, thinking they do is like having read a lot of literary text and being aware of how it has a long history of being used in serious writing.