points by DonHopkins 3 years ago

Thank you, that's an interesting response to his quotes, too! Like the proverbial plate of shrimp, this stuff has been coming up a lot recently, so I've been researching and recombining ideas from Will's old talks about Spore and with Brian Eno, David MacKay's Dasher text input system, and I've also dug up some even older unpublished thoughts from Ed Fredkin and John Cocke about a Theory of Dreams, and how they relate to LLMs and compression.

Being able to fit a program that generates lots of content on a small floppy disk or cdrom is one goal, less important today, but the important part is that thinking of procedural content generation as decompression of noise (random, user generated, contextual, or environmental) is a useful technique even with today's virtually unlimited storage and high speed delivery.

Will speaks about compression in information theoretic and arithmetic coding terms, not just referring to standard compression algorithms like "jpeg" or "mp3", but to information encoding and decoding theory, using compression techniques for procedural generation, like LLMs.

Here's a great video of Will Wright and Brian Eno discussing generative systems and demonstrating cellular automata with Mirek's Cellebration to Brian Eno's generative music, at a talk at the Long Now Foundation,:

Will Wright and Brian Eno - Generative Systems (excepts from talk):

https://www.youtube.com/watch?v=UqzVSvqXJYg

>Game designer Will Wright and musician Brian Eno discuss the generative systems used in their respective creative works. This clip features original music by Brian Eno. Will Wright and Brian Eno on "Playing with Time." In a dazzling duet Will Wright and Brian Eno give an intense clinic on the joys and techniques of "generative" creation.

Playing with Time | Brian Eno and Will Wright (entire talk):

https://www.youtube.com/watch?v=Dfc-DQorohc

>Will Wright, creator of the video games "Sim City," "The Sims," and the forthcoming "Spore," will speak on playing with time. "Playing with Time" was given on June 26, 02006 as part of Long Now's Seminar series.

Generative Music – Brian Eno (1996) (inmotionmagazine.com):

https://inmotionmagazine.com/eno1.html

https://news.ycombinator.com/item?id=24702201

Sandspiel Studio, Brian Eno, Wave Function Collapse, visual programming, cellular automata, etc:

https://news.ycombinator.com/item?id=34561910

Here's a simple low-tech pre-LLM example that shows the equivalence of compression and procedural content generation:

Take a huge text file of HN postings, and compress it with gzip or compress or some other robust compression algorithm. The better the algorithm, the more the output will look like random noise. Then slice the compressed file in half, and replace the second half with random numbers. Then uncompress it. You'll find that at the point you sliced it, it keeps on writing out almost plausible text for a while, consisting of highly probably snippets of commonly encountered words and phrases, then goes downhill towards incoherence. It's not as coherent or confident as an LLM, but the point is to show how low the bar is for using compression for procedural content generation.

LLMs are essentially a form of compression of the world's knowledge or whatever they're trained on, not just word frequencies or pixel patterns, but also concepts and ideas.

Ed Fredkin described John Cocke's Theory of Dreams, which describes dreams as a kind of procedural content generation based on running your brain's decoder over random noise inputs:

https://news.ycombinator.com/item?id=36597206

DonHopkins 1 day ago | parent | context | favorite | on: No one cares about your dreams unless you’re a fam...

ON THE SOUL: Ed Fredkin, Unpublished Manuscript

http://www.digitalphilosophy.org/wp-content/uploads/2015/07/

>The John Cocke Theory of Dreams was told to me, on the phone, late one night back in the early 1960’s. John’s complete description was contained in a very short conversation approximately as follows:

>“Hey Ed. You know about optimal encoding, right?”

>“Yup.”

>“Say the way we remember things is using a lossy optimal encoding scheme; you’d get efficient use of memory, huh?”

>“Uh huh.” “Well the decoding could take into account recent memories and sensory inputs, like sounds being heard, right?”

>“Sure!”

>“Well, if when you’re asleep, the decoder is decoding random bits (digital noise) mixed in with a few sensory inputs and taking into account recent memories and stuff like that, the output of the decoder would be a dream; huh?”

>I was stunned.

I found a video of an excellent interview Ed Fredkin in 1990 in which he explains John Cocke's Theory of Dreams in detail, beginning at 18:39, which I'll transcribe because it's so interesting and hasn't been published elsewhere (so now people and LLMs will be able to find it and learn from it too, and dream on, or even compress it, slice it in half, add noise, and decompress it to see what happens):

Ed Fredkin Talks About John Cocke - 4 May 1990: Theory of Dreams

https://youtu.be/DLCb1UV5bzU?t=1119

>I remember one time John called me up to tell me his theory of dreams. And when he first told me this, I thought, boy there's a strange theory if I ever heard one. But I have been interested in what he told me ever since. And this has to be -- it's hard for me to tell you how long ago, but it's 20, 25 years ago.

>And I'm now convinced that his theory of dreams is correct, and it's the only correct theory of dreams. And as near as I know, I don't know anyone who knows it, other than him and whoever else he's told, like me.

>But it's a beautiful theory, and it takes into account, it's really a theory based on what might be called, I wouldn't really call it information theory, but sort of information science. The knowledge we have about information, and how things are coded, and how things can be interpreted, I mean interpreted in a sort of technical computer sense.

>What his theory was, as is typical, I believe this theory was told to me in the middle of the night. But John described the following concept. Imaging the way our memories work is that they're efficient. What's known from information theory is that if you have taken some set of information and attempted to encode it in the most efficient way, then if you have succeeded, then the bits you get from your encoding scheme are indistinguishable from a random sequence of bits.

>And the reason for that is very simple thing: if the bits came out all like this: 1 1 1 1 1 0 0 0 0 0 1 1 1 1 1, then obviously it could have been encoded much more efficiently. You could say there's five 1's in a row, then five... You know, in other words, if there's all kinds of patterns to it, then it can be reduced in size by being further encoded.

>So this is a hallmark from information science of something that has been well encoded, compressed, or condensed. If the brain worked efficiently, then the following thing is true: That the things that go into our memory, if you could look into them with some kind of magic magnifying glass, like the developer they put on magnetic tape to see the actual magnetic signals, would look random.

>Ok, so that's an interesting thought. That's just applying the ideas from information theory and computer science to what might go into your brain.

>But then what John did is he took sort of an amazing leap, and asked the reverse question, which was: If you took a truly random sequence, and fed it into the decoder, what would you get?

>This is a very interesting question, because what that says is say I remember an experience I had. The experience was that I went somewhere, I went on a vacation, I went to a lake, I got a sailboat, I sailed around, there was a thunderstorm, I came back. Say I have some kind of thing like that.

>This is all compacted into these random bits. When they're interpreted, the mind has a decoding scheme, is the idea. So to be efficient, it would have to refer to a logical sequence of things. By that I mean, if it says "I went to a lake, and then I did..." Well what should happen is, the choices are: you went swimming, you went boating, you went sailing, you know.

>There's only a small number of choices, it's not an infinite number. So if you only have a small number like 5, 1 through 5, then 1 might mean I went sailing, 2 might... so on.

>So those things have to do with what is possible, and what's possible for you, and has to do with your other memories, and so on.

>So suddenly John asks the question: What if I fed in a truly random sequence of bits to the decoder? What would you get?

>Well the answers is: you would get a completely plausible sequence of events, since each choice as interpreting them is only selected from plausible events that have to do with you, but it wouldn't correspond to anything in the big story.

>And in fact, it would manufacture a dream!

>So the ideas is that if the source of the information is random noise, but it's fed into the memory decoding mechanism that works efficiently, then what you get exactly matches what a dream is!

>I think this is a very significant discovery of his, and in other words, to me it's about as important as any psychological theory I've ever seen. It makes all of the Freudian analysis of dreams fall into a totally different perspective, where what you learn is what were the categories that might have been selected that have to do with personality. But which ones were ends up just being some random noise, or something like that.

>So to me it fits in with every aspect of dreams. For instance, when you're dreaming, and something happens in the real world, like the telephone is ringing. This stimulus is often encoded in your dream. You dream that there's a phone ringing, and so on and so forth.

>Well, that works perfectly with this scheme, because if there's a phone ringing in your ears hearing it, then it becomes one of the logical things to be tapped by whatever number comes up. The only question is which context it is.

>If there is a phone ringing, the next thought has to be there's a phone ringing. But the random number tells you there's a phone ringing, and you're going to ignore it because someone else is going to answer, or various thing like that.

>So my own -- having thought about this for the maybe 20 years since I heard about, I've concluded that this is the best work on the subject that's been done on the subject that's been done so far by anyone. And I know it's not published or anything, and I'm convinced that it's a great thing. Someone ought to write it up. Or John should. That's the theory of dreams.

I posted more about Ed Fredkin recently in the discussion of his recent passing:

https://news.ycombinator.com/item?id=36429420

https://en.wikipedia.org/wiki/Edward_Fredkin

Here's another great example of applying information theory and compression techniques to efficient text entry, called "Dasher", invented by David MacKay -- think of Dasher as extremely efficiently decompressing cursor motion (or other device inputs) into text:

Dasher: information-efficient text entry

https://www.youtube.com/watch?v=ie9Se7FneXE

>Google Tech Talks, April 19, 2007

>ABSTRACT: Keyboards are inefficient for two reasons: they do not exploit the redundancy in normal language; and they waste the fine analogue capabilities of the user's motor system (fingers and eyes, for example). I describe a system intended to rectify both these inefficiencies. Dasher is a text-entry system in which a language model plays an integral role, and it's driven by continuous gestures. Users can achieve single-finger writing speeds of 35 words per minute and hands-free writing speeds of 25 words per minute. Dasher is free software, and it works in all languages, and on many platforms. Dasher is part of Debian, and there's even a little java version for your web-browser.

http://www.dasher.org.uk/

Finally, here's an interesting HN discussion about LLMs as data compression, and some interesting replies:

Ask HN: What are the data compression characteristics of LLMs?

https://news.ycombinator.com/item?id=35130027

Ask HN: DietaryNonsense 3 months ago | hide | past | favorite | 3 comments

Disclaimer: I have only shallow knowledge of LLMs and machine learning algorithms and architecture in general.

Once a model has been trained, the totality of it's knowledge is presumably encoded in it's weights, architecture, hyper-parameters, and so on. The size of all of this presumably being measurable in terms of number of bits. Accepting that the total "useful information" encoded may come with caveats about how to effectively query the model, in principal it seems like we can measure the amount of useful information that's encoded and retrievable from the model.

I do sense a challenge in equating the "raw" and "useful" forms of information in this context. An English, text-only wikipedia article about "Shitake Mushrooms" may be 30kb but we could imagine that not all of that needs to be encoded in an LLM that accurately encodes the "useful information" about Shitake mushrooms. The LLM might be able to reproduce all the facts about Shitakes that the article contained but not be able to reproduce the article itself. So in some ontologically sensitive way, the LLM performs a lossy transformation during the learning and encoding process.

I'm wondering what we know about the data storage characteristics of the useful information encoded by a given model. Is there a way in which we can measure or estimate the amount of useful information encoded by a LLM? If some LLM is trained on Wikipedia, what is the relationship between the amount of useful information it can reliably reproduce versus the size of the model relative to the source material?

In the case of the model being substantially larger than the source, can I feel metaphorically justified in likening the model to being both "tables and indices"? If the model is smaller than the source, can I feel justified in wrapping the whole operation in a "this is fancy compression" metaphor?

jeremysalwen 3 months ago | next [–]

Generative models (like LLMs) that assign probabilities to pieces of data are equivalent to compression algorithms.

To convert a generative model into a compression algorithm, you just use arithmetic coding: https://en.wikipedia.org/wiki/Arithmetic_coding.

To convert a compression algorithm into a generative model, you assign a probability to each piece of data according to the size of its compressed representation.

See also the Hutter Prize and associated FAQ: http://prize.hutter1.net/

If you wanted to specifically measure the "useful" information, you would need to have some way of sampling from the set of possible articles that contain the same "useful" information, but vary in the "useless" information, and vice versa. I think you would find that it would be difficult for you to define what the boundary is, but if you made some arbitrary choice, you could measure what you are looking for through the LLM probabilities.

PaulHoule 3 months ago | prev | next [–]

See

https://en.wikipedia.org/wiki/Hutter_Prize

GPT-3 is said to have 175 billion parameters, if those are float32s (I bet they could get away with less than that) it would be 700 GB of data. It's also said in Wikipedia that "60% percent of the weighted pre-training dataset for GPT-3 comes from a filtered version of Common Crawl consisting of 410 billion byte-pair-encoded tokens"

That would be about 680B tokens, say the average token is 5 characters, that is 3400B characters of text, such that the output is "compressed" to 20% of the input, which state-of-the-art text compressors can accomplish.

Now my figures could be off, namely they might be coding the parameters more efficiently and the average token could be longer. But it seems to make sense that if you trained a model to capture as much information as you could possibly capture out of the text it would be that size. Given that that kind of model seems to be able to spit out what it was trained on (though sometimes garbled) that might be about right.

wmf 3 months ago | prev [–]

NNCP: Lossless Data Compression with Neural Networks:

https://bellard.org/nncp/

----

Procedural Content Generation: An Overview: Gillian Smith:

http://www.gameaipro.com/GameAIPro2/GameAIPro2_Chapter40_Pro...

>One of the first examples of PCG was in the game Elite [Braben 84], where entire galaxies were generated by the computer so that there could be an expansive universe for players to explore without running afoul of memory requirements. However, unlike most modern games that incorporate PCG, Elite’s content generation was entirely deterministic, allowing the designers to have complete control over the resulting experience. In other words, Elite is really a game where PCG is used as a form of data compression. This tradition is continued in demoscenes, such as .kkrieger [.theprodukkt 04], which have the goal of maximizing the complexity of interactive scenes with a minimal code footprint. However, this is no longer the major goal for PCG systems.

>Regardless of whether creation is deterministic, one of the main tensions when creating a game with PCG is retaining some amount of control over the final product. It can be tempting to use PCG in a game because of a desire to reduce the authoring burden or make up for missing expertise—for example, a small indie team wanting to make a game with a massive world may choose to use PCG to avoid needing to painstakingly hand-author that world. But while it is relatively simple to create a system that can generate highly varied content, the challenge comes in ensuring the content’s quality and ability to meet the needs of the game.

ChainOfFools 3 years ago

I want to add this one line that tattooed itself on to my brain an early age, when I read a book about the screenwriting trade by William Goldman, of Princess Bride fame.

He cited it as the line that likewise instigated an irreversible tectonic shift in his own approach to the craft of writing. It was simply this: the realization that "poetry is compression."

I think this insight has stood up to the rise of technological leverage and may generalize extraordinarily well into some sort of philosophical ground truth, on the suspicion that given the mismatch between the scale of the universe and the two-odd kilogram lump of tissue in our skull, all of human knowledge is essentially an exercise in curated compression.