Discussion about this post

User's avatar
David Holmer's avatar

Excellent article!

One quibble: while Wikipedia is currently an excellent set of human generated summaries, it is by definition NOT a “primary source” and is a “secondary source” instead.

This distinction I think is important past pedantry as a well written Wikipedia article ALSO has links to primary sources as well. This is important as it seems like good training of AI to produce summaries would include them reading complete primary sources and then producing Wikipedia like output rather than just memorizing Wikipedia directly.

I think AI output watermarking and AI detection like pangram can also help a lot for curation of AI training data sets.

However this only removes the first order feedback loop. If AI text starts dominating what humans read on the internet then that will start influencing how humans themselves write, which WILL be picked up in future training. However I don’t see this as necessarily a bad thing. Culture evolves over time even with only humans without AI, so if AI-isms like claudish for example are accepted by large numbers of humans such that LLM compression picks it up in the training set, that’s just the new form of cultural evolution.

No posts

Ready for more?