What if the world had been flat, but infinite?

This is my post for day 24 of the Inkhaven writing retreat.

Imagine you’re in the ancestral environment. (People do that, right? That’s normal?) The only thing you know about the shape of the world is what you’ve seen with your own eyes. You and your tribe have roamed around quite a bit, so you’ve seen meadows, rivers, and mountains, and you know the earth can change a lot. But there’s always a horizon. If you actually sat down on a log and thought about it, what would the options be?

I feel like I’ve heard a lot of cultural myths about the sky. About what would happen if you somehow went up and up and up. About how the gods live there or whatever. It also seems pretty common for people to interpret the sky as a dome. Thinking of it as a dome gives you some sense that the edges might be walls or barriers of some kind, and I guess that might lead you to think of the surface of the earth as a disk of a fixed size.

Presumably one reason people think of the sky as a dome is because they can see the stars rotate smoothly. And, of course, they rotate smoothly because the earth is a sphere rotating in space, and the stars are approximately fixed and infinitely far away.

But does this actually look different than if the earth was flat, and the stars infinitely far away, and rotating around the north star? This is kind of hard to intuit, but my sense is that it would actually look the same. So why does it look so much like a dome?

Oh, right — the sun and the moon also seem to travel on the dome. Since you can see their diameter, you can tell they’re not infinitely far away, and so if they were traveling in straight lines, you’ll see them get smaller on the horizon. So maybe the sun and moon are the main reason people thought of the sky as a dome.

Let’s put that aside, and assume we have no reason to think the sky is a dome. What would people think about the horizon? If you keep walking for days on end, and new, random landscapes keep coming over the horizon, I feel like it would be very reasonable to assume that it just went on forever. (The oceans are compelling evidence for an end, but if land can suddenly turn into ocean then it stands to reason that ocean can suddenly turn into land.) And as you walked around, you would occasionally find more people. And the more you walked the more people you would find. I think it would be reasonable to conclude that the earth is filled with infinitely many people!

As a small digression, some people think that Occam’s razor means that each postulated physical object requires evidence; that the burden grows with the number of objects. They would say that, maybe seeing a few trees is evidence for a few more unseen trees, or that seeing a thousand stars is evidence for around a thousand more unseen stars, but surely infinitely many trees or stars is infinitely unlikely.

I think this is reasonable but wrong. Nature is not out there manufacturing new objects one at a time, using up some kind of finite resources, like you or I would be. Nature just is a way. And being the way of “filled with thirty trees” is not tremendously more likely than being the way of “filled with endless trees”. My formal stance on this is based on descriptive complexity; longer description length means lower probability. Since you can formally well-define “infinity” with a very small number of logical symbols, that to me makes it actually quite likely. In contrast, most specific finite numbers take a long time to describe. So the more objects you see, the more likely there are to be infinitely many.

Going back to the question of how many people there are, can you imagine how crazy the world would feel if you believed there were infinitely many people? I have grown up always knowing that the earth was finite and so were its people. It’s obviously huge and detailed, but in some sense, there’s a comprehensible limit. You can look at globes and see all the continents. You can start to memorize the names of countries, and the layout of the largest cities. But if an infinite earth was roughly evenly populated forever and ever, you would never know what might be about to come over the horizon. You could find a utopian city of gold, or a nest of dragons. If some army or plague was spreading across the land, you would just not be able to do anything about it. Presumably this is in fact how many people felt throughout history, or still feel.

There’s one more direction we could speculate about; down. This seems to be far less popular. But like, really truly, if you knew nothing about physics or the nature of matter, and all you had was a few decades of experience running around on the earth — what would you think was down there? What would you expect to happen if you just… started digging?

The sky has the really weird property that you can see right through it. And it kind of seems easier to go up. You can throw stuff pretty high. You can go climb a mountain. You can see that birds and clouds got up there somehow. But the ground is way more unforgiving, if you try to go down. Holes collapse pretty easily, so you’d have to dig out a huge cone.

I think it would be reasonable to believe that it’s just dirt and rock for miles and miles, way more than you could ever practically dig. But there would still be a fact of the matter about what’s below that. Would you believe in infinite rock? Would you believe that space itself stops? Sure, you could believe in hell, but hell has a floor. What’s under the floor? One has to wonder what a Solomonoff inductor predicts.

I’d love to read speculative fiction where some prehistoric rationalists use their spare resources to pursue a deep understanding of the universe by digging a really deep hole. And where, in that universe, the earth is an infinite flat plane, and the answer to what lies below is not the same as in our world.

Wait, IS oracle bone script older than bronze script? A mini research quest

This is my post for day 23 of the Inkhaven writing retreat.

Oracle bone script is considered the oldest known form of written Chinese. When you look at it, this makes sense; it looks more “primitive”, more pictorial, less unified in its stroke patterns. Bronze script, the next one, looks more like modern Chinese.

But hang on — oracle bone script is from 1250 BCE, and bronze script from 1200 BCE. Is fifty years really enough for a script to have evolved? In what way is oracle bone script meaningfully older?

I encountered this confusing pair of facts when first nerding out about the history of writing systems. I also mentioned it as an aside in my first Inkhaven post. For today’s post, I’m going to sit down and “livestream” my process of exploring this deeper. This means the flow and narrative of this post won’t be particularly crafted, but it will match what actually happens.

My motivation is to resolve my confusion about this particular fact, but I think that what you, the reader, should take away from this post is a detailed sense of what it’s like to do research on the modern internet. Trustworthiness is complicated, and getting answers takes a long time. The good news is that you have vast amounts of information available to you, so if you’re willing to stick with it, you can often figure out the answers to pretty niche questions.

Hypotheses

Before looking stuff up, I want to brainstorm some possibilities of how my confusion might resolve. All these ideas are just me thinking out loud; take nothing here as authoritative. Because I did read a lot about oracle bone script several months ago, my hypothesis generation below is influenced by whatever implicit knowledge I retained from then. This is also a recurring theme in research; you will have formed important but vague impressions during exploratory research, and you have to work with that when you try to nail things down later.

Okay, here are some hypotheses.

It could be that the 1250 BCE and 1200 BCE dates are representative of the oldest dated artifacts, but that historians have good reason to conclude that the two scripts have undocumented histories of different lengths. For example, maybe oracle bone script was first, but continued to be used long past the time that other forms of the script developed. This is highly precedented (postcedented?) — Egyptian hieroglyphs were calcified from a very early date, and stayed that way because it was considered the sacred form of writing. Other scripts developed alongside for practical purposes, and for smoother writing by ink.1 Written Sumerian was taught to students of cuneiform for hundreds of years after the spoken Sumerian fell out of use. Latin has the same story. Ye Olde English2 persists, especially via the influence of the KJV Bible and Shakespeare.

One possible contributing factor is that it’s harder to write by carving onto bone than other media of writing. So maybe oracle bone script stuck around for longer specifically for writing on the oracle bones.

It could be that linguists can tell that oracle bone script is older based on analyzing the etymology or shape changes of the characters. Writing systems are subject to evolutionary forces that have clear signatures. I’m not very familiar with the details, but I could find examples convincing.

It could be that some of the ancient documents explicitly say that oracle bone script is older than bronze script. This is a very unlikely hypothesis. From everything I’ve read, 100% of oracle bone script is written on oracle bones, which are all recording divination questions and results. I also believe that all of the early bronze script is on ceremonial bronze vessels, which just give names or very short inscriptions relevant to the vessel. Any ancient Chinese writing about the writing systems themselves would be (I’m guessing, but quite certain) hundreds of years older than these, and thus not particularly reliable.

Alternatively, it could be that historians are wrong.

This would be a pretty bold claim for me, a random enthusiast, to make, and I wouldn’t confidently make it without a lot more research. But part of developing an understanding of individual and collective knowledge finding processes is coming to realize just how unreliable the most reliable sources are.

History and archaeology in particular have an intense record of being confidently wrong. For a long time, the default belief was that the ruins in the Yucatan peninsula couldn’t have possibly been made by the ancestors of the native people. Top archaeologists would also do extremely dumb and destructive things, like using dynamite to excavate sites. People are systematically driven by things like status and preserving tradition. So the priors on expert consensus about oracle bone script versus bronze script just being wrong are not a rounding error.

There’s no question of the dating of the artifacts; they name kings whose reigns have known dates, and the bones have been radiocarbon dated.

Verifying the basic claims

I pulled the dates in the introduction from my head. I’d first like to triple-check them, and ideally get images of the oldest instances of the two scripts for comparison. I’d also like to find some examples of the claims that oracle bone script is older.

I’ll start by skimming through the wikipedia pages. Of course, I did this originally, but it’s fast and is a good way to find primary sources. I’ll go through them and paste relevant quotes below.

As a side note, when reading about Chinese history, the author will often reference the dynasty instead of giving a date range. The relevant ones for us here are the Shang (~1600 BCE – 1046 BCE) and the Zhou (1046 BCE – 771 BCE). The 1046 BCE date was the battle that cause the dynastic transition.

The wikipedia page for “Oracle bone script”

The info box says, “Period: c. 1250 – c. 1050 BC”.

The introduction says “Oracle bone script is the oldest attested form of written Chinese, dating to the late 2nd millennium BC.”

It then goes on to say, “The oracle bone inscriptions—along with several roughly contemporaneous bronzeware inscriptions using a different style—constitute the earliest corpus of Chinese writing, and are the direct ancestor of the Chinese family of scripts developed over the next three millennia.” This gives the citation “Boltz, William G. (1994). The Origin and Early Development of the Chinese Writing System”. Perhaps reading this book would simply answer my question, but we’ll keep going with a breadth-first search.

Later the page says, “It is generally agreed that the tradition of writing represented by oracle bone script existed prior to the first known examples, due to the attested script’s mature state. Many characters had already undergone extensive simplifications and linearizations, and techniques of semantic extension and phonetic loaning had also clearly been used by authors for some time, perhaps centuries.” That sounds useful to me, though I don’t know what many of those terms mean. It’s not clear to me whether this statement is evidence for or against oracle bone script representing an earlier form than bronze script.

Under the “Style” section, it says;3 “Along with the contemporary bronzeware script, the oracle bone script of the Late Shang period appears pictographic. The earliest oracle bone script appears even more so than examples from late in the period (thus some evolution did occur over the roughly 200-year period). Comparing the oracle bone script to both Shang and early Western Zhou period writing on bronzes, the oracle bone script is clearly greatly simplified, and rounded forms are often converted to rectilinear ones; this is thought to be due to the difficulty of engraving the bone’s hard surface, compared with the ease of writing them in the wet clay of the molds the bronzes were cast from.” That’s pretty confusing, because it makes it sound like oracle bone script is simplified from bronze script. I can easily imagine the causality going the other way, where bronze makers started smoothing out the sharp curves, given their easier medium. Overall the grammar of this paragraph manages to be impressively non-committal about the direction of causality. There is also a citation here; “Qiu Xigui (1988). Chinese Writing”. Perhaps reading this book would resolve the ambiguities.

The rest of the page continues to emphasize that oracle bone script was a full writing system, meaning that it had hundreds of years of evolution before the earlier artifact. The page also continues to cite the books by Boltz and Qiu several times.

The wikipedia page for “Chinese bronze inscriptions”

This page does not give a date range for the script above the fold. The first sentence does say (abbreviated) “Chinese bronze inscriptions … comprise Chinese writing made in several styles on ritual bronzes mainly during the Late Shang dynasty and Western Zhou dynasty”. So the date range is within the range of those dynasties, i.e. 1250 BCE – 771 BCE, but it doesn’t say how much within.

It’s also worth mentioning here that I’ve been saying “bronze script” as if it’s one thing, but as the quote above alludes to, “bronze script” refers to a category of scripts; we’re only interested in the earliest one. Bronzes with writing on them continued to be made (and preserved) up to the present day, whereas the oracle bone record stops relatively abruptly.

Continuing, the page’s introduction says, “The bronze inscriptions are one of the earliest scripts in the Chinese family of scripts, preceded by the oracle bone script.”

Later it says, “… bamboo books, which are believed to have been the main medium for writing in the Shang and Zhou dynasties. The very narrow, vertical bamboo slats of these books were not suitable for writing wide characters, and so a number of graphs were rotated 90 degrees; this style then carried over to the Shang and Zhou oracle bones and bronzes.” That sure sounds to me like the oracle bone script is not substantively older than the bronze script, and that they instead both came from the earlier bamboo-based script.

Then this page reiterates what we heard above; “The soft clay of the piece-molds used to produce the Shang to early Zhou bronzes was suitable for preserving most of the complexity of the brush-written characters on such books and other media, whereas the hard, bony surface of the oracle bones was difficult to engrave, spurring significant simplification and conversion to rectilinearity. Furthermore, some of the characters on the Shang bronzes may have been more complex than normal due to particularly conservative usage in this ritual medium…” This really sounds to me like oracle bone script is not older!

Other tertiary sources

Now that I’ve pulled all those non-committal and potentially conflicting quotes from wikipedia, the reader may be doubting my claim that oracle bone script is typically asserted to be older. To check that, I’ll just google various related terms, and tell you what some of the results say. I’ll also try to extract dates for the two scripts from the search results that seem remotely reputable (since wikipedia didn’t really give us a date for bronze script).

Okay, I did that search… and it was a pretty trash experience, epistemically speaking. One pattern I noticed is that sources were constantly conflating the dynasty range (whether Shang or “late Shang”) with the date range for the inscriptions.

Here are a few examples from sources with some degree of repute.

The Britannica article says that “The earliest known inscriptions” were the oracle bones. It also says “By 1400 BCE the script included some 2,500 to 3,000 characters”, which I think is not a date we have evidence for. It then says, “Later stages in the development of Chinese writing include the guwen…” which links to a page that sounds like “Guwen” is the bronze script.

This article on a popular Chinese language learning site claims that “The earliest Chinese characters were created using pictures or pictographs, which were originally inscribed on clay pottery and bone and then later on bronze and other metals.” and later “Bronze Writing evolved from Oracle Bone Script.” It does not mention bamboo writing.

This page from a Rutger’s professor seems to have pretty good SEO, because it was on the first page for multiple of my search terms. Reading down it, I see several yellow flags in the form of claims that contradict most of what I’ve read elsewhere. It also seems like said professor is primarily a chemist. He cites “Bronze writing” as 1400 BCE to 700 BCE, but also cites “Oracle-bone writing” as 1600 BCE 1100 BCE, which, again, I don’t think are evidenced dates.

This random page from the Harvard Art Museum says “Inscriptions cast into Shang bronze ritual vessels are among the earliest extant examples of Chinese writing.” and “Aside from bronze inscriptions, oracle bones … are the only other extant evidence of writing practice from the Shang dynasty.” This reads to me as not making a claim about which of the two are older.

Lastly, I recently went to the British museum, which had an oracle bone inscription on display. Here’s what the sign said;

Again, no mention of bronze scripts.

I’m pretty weirded out by the fact that I have failed to find a specific date for the earliest bronze inscriptions. You would think that this would be a pretty clear question to get an answer to. The description on the cover of “A Source Book of Ancient Chinese Bronze Inscriptions (Cook & Goldin 2020)” says that it “…offers English translations and commentary on over eighty-two important bronze inscriptions, ranging in date from approximately 1200 B.C.E. to 200 C.E.” So that’s something, but I was unable to find a copy of this book.

From all this and other reading, it certainly sounds like bronze scripts are no earlier than the oracle bones.

Secondary sources

Now I’m going to try moving on to what I would call secondary sources. By this I mean academic books written by experts in the field.

The Ancestral Landscape (Keightley 2000)

In my past research, I’ve previously read the entirety of “The Ancestral Landscape: Time, Space, and Community in Late Shang China (ca. 1200–1045 B.C.)” by David N. Keightley. As far as I can tell, he is considered a reputable, and a leading (Western) researcher in the field.

Skimming through it again now, I realized that this book is mostly about what conclusions we can draw about the Shang people and culture using the oracle bone inscriptions as our source of information. It doesn’t talk that much about the development of the script, though it does talk a fair bit about the script itself. Unfortunately the word “bronze” does not appear in the index.

Does Keightley claim that oracle bone script is the oldest? Here’s a sentence I found in the preface; “These records, the earliest body of writing yet found in East Asia, were produced in the following way.” That actually sounds fair to me. Claiming that the oracle bones are the earliest body of writing is much different than claiming that bronze script descended from oracle bone script. I haven’t quite figured out how big the corpus of (late Shang) bronze script is yet, but it sounds plausibly small enough not to constitute a “body”.

Chinese Writing (Qiu 2000)

Now let’s check out one of the books heavily citied by the two wikipedia articles. https://starlingdb.org/Texts/Students/Qiu%20Xigui/Chinese%20Writing%20%282000%29.pdf This book is a 2000 English translation of the 1988 Chinese original. The preface makes it sound like it’s also been quite updated, so perhaps we can consider the information to be up to date circa 2000.

From the table of contents, there could be a lot of relevant sections, but it does not seem to contain the words “oracle” or “bronze”. (This PDF is not OCR’d so I can’t use the find function.) The index contains several page numbers for both scripts (which tend to have overlapping ranges). From reading through all these pages, it sounds to me like Qiu treats the two scrips as on-par with each other.

From page 29; “The earliest relatively substantial examples of ancient Chinese writing discovered so far are the bone and bronze inscriptions of the late Shang dynasty.”

From page 63; “It should be pointed out first that in terms of their structure, bone and bronze graphs exhibit different characteristics. During the Shang period the writing brush was the primary writing implement in use. … Graphs appearing on bronzes retain the features of brush-written characters, whereas those written on bone do not. As the Shang rulers frequently made divinations, the number of divinatory notations that had to be inscribed on bones was quite large. Inscribing characters on a medium as hard as bone is a time-consuming and strenuous task. For the sake of efficiency, engravers out of necessity altered the forms of the brush-written characters… Bone script can be viewed as a rather peculiar form of the popular script of that era, whereas the bronze script of that period for the most part may be viewed as a formal script.”

So without going deeper and reading the whole book, my judgement is that this source does not support the idea the oracle bone script precedes bronze script, and instead supports the idea that they were contemporary scripts adapted to their function and medium.

Overall, this book seems extremely reasonable and balanced.

The Origin and Early Development of the Chinese Writing System (Boltz 1994)

Reading through the index, it says almost nothing about bronzes. It seems to mostly be an academic defense against the (then) popular idea that Chinese was pictographic. I now realize that while this book was heavily cited by the wikipedia page for oracle bone script, it was not cited on the page for Chinese bronze inscriptions. Oops! It does not seem to make a strong claim about oracle bone script in particular being the predecessor of all other Chinese scripts.

Conclusion

My conclusion is that oracle bone script is not a predecessor to bronze script. Instead, they both existed at the same time as each other and as a more common script that has not survived due to being written on perishable materials like bamboo.

However, my conclusion is also not that “historians were wrong” — it’s that all the tertiary sources are wrong. This should have been an obvious hypothesis, but I failed to think of it.

So what happened? Why does everyone say oracle bone script is the older predecessor? I think there are several contributing factors.

  • Oracle bone script “feels” older. Both because it looks more primitive and because bones feel older than bronze.
  • It may be that technically, the oldest known date for an oracle bone inscription is slightly older than the oldest known date for a bronze inscription.
  • The dates of all the known oracle bones are tightly clustered in the 200 year range which also happens to be the beginning of the date range for bronze inscriptions. Since oracle bone script “died out” much earlier than the bronze script, that kinda makes it feel older.
  • Scholarship in early Chinese writing is niche, and mostly not translated into English.

Though this post is very long, it was only one day of research, and again, I am by no means an expert! So if you are, or happen to know key facts I missed, please let me know!

And if you’re a random bystander who’s considering going into research; hopefully this little adventure helped you decide. I honestly enjoyed it a lot.

  1. An Egyptologist will tell you that hieroglyphs changed a ton over the millennia of use; this is true, but we’re talking about differences of a different magnitude. ↩︎
  2. Actually Early Modern English. Old English is what Beowulf was written in, and is totally unreadable to a modern English reader. ↩︎
  3. Throughout this post, all emphasis in quotes is mine. ↩︎

Knowing thyself does not imply fixing thyself

This is my post for day 22 of the Inkhaven writing retreat.

The basic promise of science is this; by using the scientific method, you can figure out how the world works, and thereby take better actions to get the outcomes you want. This is the primary justification for funding scientific endeavors with taxpayer dollars. There is of course also the justification that scientific inquiry has value in its own right. But this isn’t valued by everyone, and is overall a much harder sell.

In some fields, like astronomy, it’s clear that the main activity is observation and not manipulation. No one is expecting astronomers to figure out how to make the sun be less bright to reduce global warming. But I think a lot of people are still expecting that studying the deep nature of neutron stars gives us a better chance of discovering some kind of double-hyper-fusion, or something.

And people often talk about the danger of attempting to apply the knowledge gained from investigating nature. Whether it’s nuclear bombs or Jurassic park, the harms of science-fueled technological development have been vividly impressed into the public’s mind.

In contrast, I almost never hear anyone talk about the null outcome that some domains of science are substantially inactionable.

Investigating the mind is one such category with dubious applicability. Putting aside the hard problem of consciousness, I would claim that we still have basically no idea how “thinking” works. We have a lot of data about neurons and synapses, and also a lot of fMRI data, the value of which is very unclear. We have no shortage of ideas about how thinking could work. And at this point, we’ve literally built a brand new intelligence (using methods that specifically obfuscate the entire structure). But we still don’t know the basic, architectural principles behind the human mind.

Personally, I have some kind of moderately severe problem with the way my mind works, which can be expressed via attention, motivation or executive function. I spent a lot of my 20s trying to understand what was going on at the level of psychological investigation. Did I have false subconscious beliefs? Did I have underlying values that I wasn’t aware of? Did anything in my childhood cause this? And I figured out a lot of stuff! I outlined several models of my psychological content which resonated and were consistent with other facts about my life. But change was lagging.

Sometimes, when you do enough introspection to excavate an underlying belief, the realization of it causes the relevant problem to almost autonomously resolve itself. Perhaps you fantasized about getting a puppy but were conflicted about it making your house too messy. Through introspection you realize that you mostly wanted a puppy to cure your loneliness, and that you wanted a clean house to feel a sense of control. And now that you’ve realized this, it feels clear that a better solution to both of these is to put more focus on finding a romantic partner who would enthusiastically help you gain better skills to control your life. But my big problem was not resolving itself despite my “discovery” of some underlying content that seemed very related.

Another way to fix some motivation or executive function problems is to set up strong habits. Habits are an extremely real, extremely reliable neural/psychological mechanism that humanity has written about since Aristotle.

There are several self-help books that describe the phenomenon of habits in great detail. The habit will have a contextual cue or trigger. This trigger will cause you to take the habitual action. The action will then result in you experiencing a reward. This reward reinforces the trigger-action pair, making the habit more likely in the future.

These books are written in a frame of problem-solving; they have little diagrams or flow charts that help you figure out how to install or uninstall a given good or bad habit. But as I read through these books, brainstorming for hours and trying countless little modifications to my life, I found that nothing really stuck. The books are supported by a mountain of habit science, but I do not believe we have an equivalently supportive habit engineering.

Certainly some attempts to deliberately install habits work. Many of the habits you already have are from you starting to take actions that you thought would have positive effects, and being reinforced by that. But those were also likely easy enough that you didn’t need a whole habits framework to do it. Perhaps there is something like an efficient market for habits, where the habits that are worth the effort of installing are ones that you’ve already installed.

I think that failing to acknowledge this science vs. engineering or observational vs. interventional distinction is a generalized sin for self-help books, particularly the ones that claim to be based on science. It’s not an optimistic piece of advice, but keeping it in mind and tracking them separately can save you a huge amount of time.

Inkhaven check-in: how is blogging going?

It’s day 21 of 30 days of daily blogging at the Inkhaven writing retreat.

The social accountability structure is doing an excellent job at causing me to actually write something every day. This kind of structure is rare and valuable; you can’t just go out and get 40 people to do a daily goal with you and mutually support each other. I’m glad I took advantage of it. I’m also quite glad that I’ll have a little portfolio of 30 blog posts, like a kind of souvenir. Probably the most useful thing I’ll have gotten out of it is the information of what it was like when I tried to write every day. I’m getting lots of rapid data about how easy it is, what topics I end up choosing, and how my writing responds to a forceful incentive to publish faster. Even if I stop daily blogging, I can use this bundle of experience as useful data when I want to plan future writing, or reflect on my relationship to writing.

That said, I think that by day 20 the marginal information accumulation has kind of petered out. I don’t feel strained for ideas or anything, so I’m pretty sure the next ten days could just go exactly like the last ten.

Given that, I’ve been experiencing an increasingly strong sense of “I don’t wanna” in relation to the task of writing a post each day. Now, this feeling on its own doesn’t necessarily mean much. It’s pretty common for highly skilled or productive individuals to report that their hesitation to wake up early, do the hard work, get on stage again, etc. never goes away, and the key to succeeding was to build a system (internally or externally) such that they do it anyway, despite that feeling. So it would be kinda lame if I let this sense of “I don’t wanna write a post today” affect me in and of itself.

But when I look into that feeling a bit, part of what’s happening is that the writing doesn’t feel as valuable as I thought it might be.

A lot of how I’ve been writing daily posts is by just lowering my standards. I haven’t written anything that I’m embarrassed about, and it’s not like I’m dumping stream-of-consciousness morning pages onto the internet. But the posts are… just fine? They’re essentially all notably worse versions of what I “could” have written if I’d spent more time on them. I would like them to be better fact-checked, and illustrated, and with much more care in the construction of the expression. I want to take more time to find the phrases that really hit with resonance. It may be that writing three of the worse posts is more valuable to readers than me spending three times as long to make the better version of one of them. But… I don’t really want to write the worse versions? That is not very satisfying to my inner craftsman.

As you can see, I’m still figuring out how I feel about this.

I’m also getting less feedback and positive reception than I expected. This could be just down to the fact that I am doing no promotion of my content whatsoever, but it’s still a negative update on the value of writing.

There’s also the bit where the activity of getting to know the other people in the Inkhaven cohort has been, thus far, a direct trade-off against writing. I’ve mostly prioritized writing (and my day job) and so I really haven’t gotten to know people. This program feels like exactly the kind where you unexpectedly life-long friends, so I’m kind of sad about having to write mediocre blog posts instead of doing that.

Your values constrain your actions

This is my post for day 20 of the Inkhaven writing retreat.

Every day, everyone wakes up all around the world and gets down to business trying to achieve their values; raising their children, producing goods, having positive experiences, and generally staying alive. Much of life has the sense of trying to push forward toward something, towards your values, and being constrained or rate-limited somehow.

It’s easy to see how your actions are constrained by things like money, skills, or social support. These constraints are like physical walls delineating the room you can act within. Inside the room is all the actions you can take immediately and freely, whether or not they effectively achieve your values. Walls can be broken down, but it takes quite a lot more effort, planning, and hard trade-offs.

It’s less natural to look at your values themselves as constraints. You “could” swerve the car into the oncoming lane, but you overwhelmingly don’t want to.

I think it can be a useful perspective to view yourself as physically constrained by your values, just as much as you are constrained by money or skills, even if it’s a constraint that you’ll never try to overcome. (In this post I’m only referring to terminal values, or what philosopher Paul Tillich called ultimate concerns, which are the things you value in and of themselves. I am not including instrumental values, which are things you only value because they lead to other values.)

Values as constraints is a pretty funny way to look at things, like putting the cart before the horse. The primary relationship between your actions and your values is that your values are why you take the actions you do take. It’s not as if your values are “take as many actions as possible”.

But it’s still kinda true, though. Scott Garrabrant once quipped that an agent was something whose type signature was (A → B) → A. That is, if the agent predicts that action A will lead to outcome B, and the agent values outcome B, it will take action A. Similarly, (A → ¬B) → ¬A.

This can also be a useful perspective for viewing other people.

I have sometimes been confused about why some of my friends seemed to struggle with certain things. It was easy to consider that they might have had less skills, or had different life experiences, or sensory sensitivities, or different brain chemistry that produces more anxiety, or something. It took me longer to realized that they simply had values that I didn’t. Once I could internalize that they really did have those different values, it was obvious that their action space was more limited, and their struggles made sense.

Or, maybe other people are missing values that you have. This would give them more options for acting. This is one reason that powerful people are more likely to be be sociopaths. I physically could not take the action of hurting people in the way that many politicians or businessmen do. But they can take that action, because (in part) they literally don’t care. All else equal, a larger action space implies a higher probability of achieving your goal. Perhaps it is tempting to think something like “curse my pro-social values, if only I didn’t have them, then I could gain great political power, and with it, do higher-leverage pro-social things”. But like. That doesn’t really make sense. As a disclaimer, I don’t mean to imply that all politicians and businessmen are sociopaths, or that society is doomed (by this particular selection effect). Just that it’s something you should have in your model of society.

This idea applies less cleanly to people whose values are less stable and coherent. A more coherent mind might value both apples and oranges, with some weighing between them. If it has to make a decision that trades off between apples and oranges, then it will just apply the weights to decide. A less coherent mind might simply contain two subsystems, one which values only apples, and one which values only oranges. This mind would also have some kind of supervising system that controls when each subsystem runs. In this case, the conflict between the values will be a genuine conflict, and one of the subsystems might figure out how to destroy the other one. This is more like how I would describe becoming corrupted.

A review of Red Heart, the new AI alignment novel

This is my post for day 19 of the Inkhaven writing retreat.

I recently read Red Heart, a spy novel taking place in the core of a Chinese AGI project. Disclaimer that the author is my friend, and that I’m ideologically incentivized to promote stuff about AI safety! That said, I think you should read it. If nothing else, it’s a fun read.

The first half of the novel feels very clean, crisp and controlled. The data center and office building are all brand new and in a remote location. As a top-secret Chinese government project, the culture of the office is very obedient. Chen Bai, our spy protagonist is constantly monitoring what he says, and the implications of what he sees. His job on the project is to ensure the value alignment of the AGI, so even is official job is to be paranoid. He has no contact with family or friends, and even his apartment is newly built for the project. He works most waking hours.

I used to be a software engineer in San Francisco and am now a researcher in AI safety, so much of the setting and content of the first half felt very normal to me. I think that for readers further from the setting, the rows of monitors, white board sessions and terminal commands could feel more novel and interesting.

At the halfway point, we really hit a different gear. Bai has been as careful as he can, and now he needs to start taking risks. We also start getting deeper perspectives from the other characters — a new friend, a boss, a love interest — which had previously been chess pieces. Inside Bai’s head is not the best place to spend a few hours.

Yunna, the AGI, feels believable to me, though that’s largely because she is very much like a human, and I believe that human-like AGIs are quite plausible. When we met Yunna she was already in a pretty coherent and generally intelligent state. I would have liked to see more of the transition between a ChatGPT-like model and the Yunna we meet. I have no complaints about any of the “sci-fi” elements being unrealistic, unlike virtually every other sci-fi media I’ve consumed.

Separate from the AI themes, I really enjoyed hearing characters speak from the perspective of a Chinese worldview. I’ve read some about the history of China, but I’ve spent essentially no time learning about the perspective of native Chinese people. The only judgement I get exposed to is to the basic “China bad” American take. In contrast I found the expressions of the characters in Red Heart quite reasonable and believable. Of course, Harms is not culturally Chinese, so I read it with that distance in mind. But everything that I spot-checked looked valid to me. Hearing an ideological character justify themselves by citing the Rectification of Names was a fun detail to investigate.

This story is one of desperation, and of well-meaning people being strained by too many constraining forces. This dynamic is happening in real life, and society is not ready for how the strains may break.

There are two obvious endings for a novel about AGI, which are “utopia” or “everyone dies”. Harms successfully navigates us into something more interesting, without undermining the main messages around AI risk.

My textbook choosing ritual

This is my post for day 18 of the Inkhaven writing retreat.

For the kind of things that I want to know, and the way that I want to know them, I find that textbooks are a pretty effective type of resource. But even starting to read a textbook is a pretty big investment for me, so I have a fairly heavy process of selection.

Or rather, that’s one reason the process is heavy. Another reason is that I love it.

Knowledge, and especially the artifacts of humanity’s quest for knowledge, are essentially religious objects for me. An entire library of them is overwhelming. So given this task of finding a textbook for a specific, endorsed purpose, I indulge my desire to worship.

I start the process when I have a fairly well defined scope of what I want to understand. Sometimes it’s relatively specific (“could someone please tell me what variational inference is”) and sometimes not (“actually I just want to read the first 100 pages of whatever an archaeology student would learn about stone tools“).

Shockingly, my first step is actually to use google. Usually I search something like “[thing] textbooks” or “best [thing] textbooks”. My goal here is not actually to find the best textbook about [thing], since there usually isn’t one. (If there is, I often find it on this step, and that saves me a lot of time.) Instead, my goal is to get a collection of a few of the most common textbooks on the topic, both to potentially check out those specific ones, and to seed the next step.

The next step is that I log onto my library’s search site. It is essential to this process that I have access to a large academic library system. I find the record for each of these books and write down the Library of Congress code. Often the codes are really close to each other, but also they’re often not, and that can tell me how long it might take to scan all the relevant areas. Some books about stone tools might be under archaeology, and others under geology. As an aside, I will say that every year it gets harder to convince the system that no, I really do want search results for only physical books, please.

Then I walk over to the relevant libraries, and bring a large bag, just in case.

For each roughly clustered section of the LOC codes, I find those specific books and pull each of them out a couple inches as a form of bookmarking them on the shelf. Then, I scan forwards and then backward to find the bounds of the contiguous section of the shelf that will plausibly contain books about my desired topic. When I do, I also pull out those first and last books a few inches. Then, I begin the process of quickly scanning every book in between, and pulling out the ones I’m interested in looking into deeper.

This part is a little bit crazy and excessive. Sometimes I will find that the section is like, five whole shelves, and then I have to give up and rescope my search. But usually I can actually scan all the books. My local academic library system is big enough to have lots of books on niche topics, but it’s not like I’m scanning through every existing textbook on the topic. And the fact that my library has a physical copy is enough of a signal of quality that I figure it’s worth scanning over. But I emphasize that this is really not normal and if you are just a random student reading this post then do not take this as advice on best practices.

During this first scan, virtually every physical aspect of the book gives me useful information. The title is obviously the most important. LOC codes puts the date of publication at the end, so I can quickly filter out very old books. Sometimes I want to rule out books that are too thick, and other times I’ll deduce that a book is too thin to be a comprehensive introduction.

I’ve also learned that, at least at my library, some styles of binding mean specific things. For example, there’s a binding that means the book was written with a typewriter. Another binding means it’s in a foreign language. Another binding usually means that the book was so popular that it wore down and they had to rebind it with a tougher binding. These are often the “classic” textbooks in the field, the ones assigned in classes, and there will usually be multiple copies.

Doing this full-scan process gives me a bunch of cool implicit information about the field as a whole. I can see how many books have certain adjectives in the title. I can see which books were the founding texts, and which books are trying to be the revised, modern editions. I can see how prolific certain authors are. I can tell which subjects were more popular with soviet mathematicians, or that the chaos theory boom led to all the 1990s dynamical systems textbooks to have “chaos” in the title. I’m probably learning lots of things that I never realize I’m learning. But also, it’s part of the ritual of worship.

After this scan I take a step back to look at how many books I’ve pulled out. Usually I do a big sigh and check the time to make sure it’s appropriate to spend another hour sitting in front of this shelf. Time does not exist during this ritual.

For the next phase, I will pull each “bookmarked” book off the shelf, and start scanning the front matter & back matter. My goal here is to decide whether this will be one of the books I take with me over to the library tables, to read in more depth. I can really only do this with six or eight books, so I have to be pretty picky. I read the back if I haven’t already. This is the first time that I hear a voice tell me about the book. There’s a pretty generic formula for what the backs of textbooks say, so it’s not too informative. But it can sometimes tell me things like whether the author wrote this book in order to convince people that their special sub-interest is important.

If I’m looking for a more specific topic like “variational inference”, then I’ll check if it’s in the table of contents or the index. The TOC will also very efficiently give me a sense of how the author thinks about the subject. Reading half a dozen TOCs about the same subject gives me a really good sense of whether the field as a whole has converged on one way to present the concepts. All the TOCs in semigroup theory are exactly the same, whereas the TOCs for functional analysis can vary substantially.

This can sometimes be enough to decide whether a book goes in the table pile, but I’m often reading the preface or introduction as well. This is the part of the ritual that really starts to pay off spiritually. Sure, sometimes the main thing I learn is that the author is out of touch with what a student should find elementary, or that the author is only publishing this book to gain social clout. The bulk of textbooks are written in a fairly detached, objective style. But sometimes I find that the author relates to the subject with the same sense of meaningfulness as I do. With their decades of expertise, they can help me begin to see the ways in which the subject reflects deeper aspects of nature.

Here’s an example from the preface of Computational Complexity by Christos H. Papadimitriou.

At the risk of burdening the reader so early with a message that will be heard rather frequently and loudly throughout the book’s twenty chapters, my point of view is this: I see complexity as the intricate and exquisite interplay between computation (complexity classes) and applications (that is, problems).

Here’s another, from The Art of Turing Computability by Robert I. Soare.

…It is not enough to state a valid theorem with a correct proof. We must see a sense of beauty in how it relates to what came before, what will come after, the definitions, why it is the right theorem, with the right proof, in the right place. …. The first aim of this book is to present the craft of computability, but the second and more important goal is to teach the reader to see the figure inside the block of marble.

I won’t always come to agree with or find use of the perspective of these authors. But if possible, I’d like to be shown the world by a certain kind of mind, the kind that feels compelled to describe its object of study as an “exquisite interplay”, or to compare it to a statue. This ritual lets me explore how different minds relate to the same ideas, and lets me find the right guide to follow on the path.

I store some memories spatially and I don’t know why

This is my post for day 17 of the Inkhaven writing retreat.

Every so often, I have this conversation:

Them: So you know how the other day we talked about whether we should leave for our trip on that sunday or monday?
Me: …doesn’t sound familiar…
Them: And you said it depended on what work you had left to do that weekend…
Me: Hm… where were we when we had the conversation?
Them: Um… we had just arrived at my house and I had started making food-
Me: Ooooh yeah yeah okay. And I was sitting on the black stool facing the clock. Okay cool, I remember the conversation now, please continue.

…What the heck is up with this? Does it happen to anyone else? Apparently, my brain decides to index conversations to be efficiently looked up by quite precisely where I was in physical space when the conversation occurred. I have no conscious experience of this indexing happening. It’s also pretty strange that it happens for locations that I use on a regular or even daily basis; it’s not like I could just start listing all the conversations I’ve had while sitting on that kitchen stool.

I do believe that I’m quite above-average aware of what’s happening in my visual field. I always notice when people come in and out of a room, I tend to see new objects or decor right away, and I somehow spot every insect. I’m often the first to spot a leak or mold. I almost never run into stuff or knock things over. So maybe it’s just increased attention to my surroundings?


Here’s a similar pattern I’ve noticed.

I’m in a phase of my life where I read a lot of books, and especially textbooks. My field of study is interdisciplinary, and I am frequently looking up something that I’ve read before. When I do, I will frequently have the sense of roughly where it was, physically, in the book. This includes:

  • how far into the book,
  • whether it’s on the left or right page,
  • how far down the page,
  • roughly where within a paragraph it is,
  • and a vague sense of what the rest of the page looks like.

To be clear, I’m not claiming that I have any kind of “photographic” memory. I have no idea what almost all of these books say. I don’t have any degree of verbatim retention. But when I remember that there was a particular interesting part and want to go look for it, my brain brings up these visuo-spatial associations. These associations feel blurry but confident, like some kind of hash function. Textbooks are heavily formatted, so there will be lots of white space, diagrams, section headers et cetera to anchor off. When I try to recall the “location” of events in flat prose fiction books, nothing comes up.

This is, I think, one reason why I have struggled to switch over to digital forms of books. I’ve tried it a lot, but they always fade out of use. There are many other reasons (if they’re not on my shelf I tend to forget it exists, I find physical books far easier to skim) but the fact that I can’t physically index my knowledge to it is noticeable. It’s just some big infinite scroll that looks and feels indistinguishable from all the other big infinite scrolls.


I’d love to hear how others relate to either of these experiences!

What would my 12-year-old self think of agent foundations?

This is my post for day 16 of the Inkhaven writing retreat.

I knew I wanted to do science and math from a very early age. And I didn’t want to spend my life investigating just some particular phenomenon; I wanted to understand “everything”. You obviously can’t do that in a literal sense, so I focused on understanding things that were increasingly general. Generalizations are in some sense more “efficient” ways of understand things.

In physics, the field that claims to see the “theory of everything”, there are two obvious directions you can go with this. One is “up”, to astronomy and cosmology and the overall structure of the universe. The other is “down”, where you can look at the smallest particles. I was very into both of these, though it’s clear that the “down” direction is in some sense more fundamental. If you understand the laws behind the behavior of the smaller things, you can, in theory, use them to calculate what will happen to the bigger things. At 12 I knew that people had made huge progress toward finding the fundamental physical laws, and I was very excited to catch up on it.

But there also seem to be some other “directions” to generalize in if you want to efficiently understand everything. Mathematics is one of them, which is something like the “symbolic” direction. Philosophy is perhaps in the “conceptual” direction. And there’s also a direction that is something like the study of yourself, of how minds work. The study of what’s up with being the kind of thing that is inside the universe, observing and trying to understand it. This is a meta type of direction.

Over the years my preferences between these has shifted, but my 12-year-old self would not have been too surprised if I ended up going deep in any of these directions. Treatise on ontology? Awesome. Unified theory of the neocortex? Let’s go. But what’s agent foundations?

Well, I moved into agent foundations because I decided we needed to solve a problem, namely existential risk from AI. But it’s mostly trying to help with that problem by figuring out what the heck is going on with the phenomenon of agents. Which is to say, by understanding it.

I think my younger self would be pretty confused for a while that I’m into this, but I could probably explain it given enough time.

Above I mentioned the study of physics at the biggest scale and the smallest scale. There is obviously a lot going on in the middle, but from some perspective it feels mostly arbitrary. Like, there just happens to be water and flowers and binary star systems. All those things are interesting insofar as they are in the set of “everything”, but it doesn’t feel like understanding them has much generalization power.

I claim (to my 12-year-old self) that there is actually a generalized theory of things going on in the middle. That is, a generalized theory, not about specifically what’s going on in the middle, but about what it means that something could be said to be going on in the middle. (That sentence may have lost the reader. It also may have lost my 12-year-old self, but he’s now very excited to understand what I meant by it.)

For example, what exactly does it mean when we say that Newtonian mechanics is a good approximation of the true laws of physics? If you handed someone only the true laws of physics, how could they have figured out, in principle, that Newtonian mechanics was a good approximation of what they were holding? Are there other possible good approximations they could have figured out instead? Can we well-define the set of all possible good approximations, given some true physical laws?

This is relevant to agent foundations because agents have models of the world inside them. They use these models to successfully achieve their goals, so the models must be good approximations by that standard. If we want the agent to achieve our goals, then it probably needs a world model that is compatible with ours, at least in the parts that describe our goals.

We currently do not know how to formally state this, and I think that’s a barrier to being able to ensure it in practice.

Another word you could use for “world model” or “good approximation” is “theory”, so in some sense this part of my work is studying the theory of theories. Which, yeah, my 12-year-old self would be pretty thrilled about.

My favorite walk

This is my post for day 14 of the Inkhaven writing retreat.

My boyfriend and I live four short blocks away from each other. An eight minute walk. We are in our eighth year. I walk there often, most often at night, around 7:30. For some part of the year, this is during sunset. Sunset is when the neighborhood cats come out. Donut lives closest to me; he’s a black cat, the rockstar of the neighborhood. A couple houses to the right is a crazy grey and white cat. It will chase you and tackle you, harass you and hiss at you. But in a playful way.

I leave my door and turn left, walking half a block toward the sun and the bay. Then I turn left again. At this house, there’s a Little Free Library box. I always stop and look in it. The books are different almost every day. I’ve wondered if the owners cycle the books. I can’t imagine passersby taking them all. I once left a copy of the Epic of Gilgamesh in the box. It was the Penguin edition from 1960. It was outdated; since then, we’ve found many more tablets, which add to our knowledge of the Epic. The book was gone the next day.

Then I go through the park. The park is very long and skinny, and I only cross it for half a block. I cross the first street. This one is reasonably low-traffic, but the visibility isn’t great. I walk down the second block. There are memories here, but they’re all faint. There were bees once. I look into the house that never closes their curtains. I rarely see them in there. I think they’ve moved out, now; the curtains have started being closed sometimes.

I cross the second street. This intersection has barriers that turn it into two separate right-angle turns. Only emergency vehicles should drive straight through. This makes it almost always empty of cars, much nicer to cross. Half the time that I encounter cars here, they’re making a mistake and turn around.

I get to the next street. On the left is the Mexican tile & ceramics store. Every item in this store is heart-achingly beautiful. I’ve bought some from there before. But I mostly don’t have need for tiles.

Crossing this street is the worst. It’s super wide. The pavement reflects the sun and it’s hotter somehow and the whole world suddenly feels like Los Angeles. There’s no stop light. There’s only the lights where you can press a button and make them blink. The blinking is not particularly persuasive.

Crossing this street is the worst, but it’s better than going the other way. The other way, I’d have to cross the worst intersection I know of. That intersection turns drivers psychotic. I’ve had to jump out of the way of the cars more than once. Those same cars are the ones that also cross this intersection, but by this point they’ve regained their sanity. That intersection has a stop light and a pedestrian cross light. Apparently, it is not particularly persuasive either.

Crossing this street is the worst, but it’s the last one.

I cross the street and pass a burger place like an off-brand McDonalds. I’ve never been in. Just the other day I realized that it’s open 24 hours a day. I don’t know anyone who goes there.

At the end of the block I look into the other house whose curtains never close. There is what looks like a stripper pole in there, but I’ve never seen anyone use it. I almost never see anyone in there. I think new people have moved in, now; I see them, and they close the curtains sometimes.

I turn left, and walk toward my boyfriend’s house. There are three ways to get into his apartment. If I go through the front door, there might be people hanging out in the living room. It’s easy to get caught up in conversation with them. But I’m not here for that; I’m here for him. My favorite way is going up the driveway and into the back door. It’s closer to his bedroom. But if I arrive after dark, my passing makes lights turn on. Click, click, click, click. Super bright. Too much attention. I’m trying to remember to go through the gate down the other side. The gate used to be closed all the time. It’s open now, but it still feels weird to walk past the downstairs unit’s door.

He’s always surprised when I show up. He never checks his phone.

Sometimes he comes right back with me to my house for the night. We call it doing a fetch. The other day, we did this just when it started downpouring. We were prepared, me with a raincoat, him with a poncho. We splish-splashed in the puddles and torrents all the way home. It was the best part of my day.