# Metaphor Hacker
> Hacking Metaphors, Frames and Other Ideas. Long-form essays by Dominik Lukeš on conceptual metaphor, cognitive linguistics, education, writing, and language technology. Originally published at https://metaphorhacker.net/ from 2010–2025.
---
# Pages
## About
URL: https://metaphorhacker.net/about/
Image via Wikipedia
Metaphors are not just something extra we use when we're feeling poetic or at a loss for le mot juste, they are all over our minds, texts and conversations. Just like conjunctions, tenses or words. And just like anything else, they can be used for good or ill, on purpose or without conscious regard. Their meanings can be exposed, explored and exorcised. They can be brought from the dead by fresh perspectives or trodden into the ground by frequent use. They may bring us into the very heights of ecstasy or they may pass by unnoticed. They elluminate and obscure, lead and mislead, bring life and death. They can be too constrained or they can be taken too far. They can be wrong and they can be right. And they can be hacked.
Hacking metaphors means taking them apart seeing how they work and putting them back together in a creative and useful way. People hack metaphors all the time without realizing what they're doing and often getting into trouble by not recognizing that this is what they're doing. Paying a bit more attention to how metaphors work and how they can be made work differently may make their hacking an easier process.
Oh, and ...
Metaphor doesn't really exist as a separate clearly delineated concept. It is really only one expression of a more general cognitive faculty I call conceptual framing (http://en.wikipedia.org/wiki/Framing_(social_sciences)). Depending on who you ask, it is different from or the same as simile, analogy, allegory and closely related or in opposition to metonymy, synechdoche (http://en.wikipedia.org/wiki/Synechdoche), irony, and a host of other tropes (http://en.wikipedia.org/wiki/Figure_of_speech#Tropes). On this site, these distinctions don't matter. All of the above rely on the same conceptual structures and metaphor is just as good a label as any for them. Sort of a metaphor, really...
## Blog on Request
URL: https://metaphorhacker.net/blog-on-request/
I needed a place to keep track of ideas I may want to blog about. So here it is - completely drafty and non-binding. If anyone has any preference for what the next post should be about, I'd be more than happy to oblige:
- Narrative fallacies of multilingual education
- There's no such thing as science
- The conceptual structure of capitalism
- Adjectives as metaphors
- Metaphorical structure of the moral compass
- Negotiation over the conceptual structure of political controversies
- Mr Data's emotion chip and the fundamental gap in our understanding of the mind
- What if the legal system was based on metaphorical rather than logical rationality (and why it in part is already)
- Appeals to reason as a rhetorical ploy
- Turing test of educational achievement: A signalling metaphor
- Cultures of life and death and metaphorical blindness
- Is intellectual property theft?
- Why don't metaphorical hawks eat metaphorical doves?: The case of the missing metaphor
- ...
---
# Posts (82 total, newest first)
## Schemas and Propositions: What and where I've been writing and what's next
Date: 2025-08-24
URL: https://metaphorhacker.net/2025/08/schemas-and-propositions-what-and-where-ive-been-writing-and-whats-next/
Categories: Extended writing
## What's been happening?
It has been a little over two years since I posted on this blog last. I have been doing a lot of writing elsewhere on similar topics - mostly Large Language Models and other AI things - it just never added up to a blog post here.
### My writing exiles
Here are some of the places where I have been writing:
- Random musings on practical epistemology:
- Promising Paragraphs (https://promisingparagraphs.substack.com/)
- X (https://x.com/techczech)
- Questions of Artificial Intelligence:
- Trends in AI (https://ai-trends.notion.site/)
- Semantic Machines (https://semanticmachines.notion.site/)
- Practical AI (https://practical-ai.notion.site/)
- Experiments in AI Creation (https://aicreationexperiments.notion.site/)
- AI in Academic Practice (https://www.linkedin.com/newsletters/ai-in-academic-practice-7045331862348554240)
- Questions of practical pedagogy:
1. Deliberate Practice as a Universal Learning Method (https://deliberatepractice.notion.site/)
2. Building Your Language Muscle (https://buildingyourlanguagemuscle.notion.site/)
3. Creating and Using Instructional Videos (https://instructionalvideos.notion.site/)
4. Principles of Readability (https://readabilityprinciples.notion.site/)
5. Digital Technologies for Academic Productivity (https://academicproductivity.notion.site/)
6. Czech Language Navigator Companion (https://czechonline.notion.site/)
7. Paragraph a Day for a Month: Academic Writing Boot Up Diary Template (https://paragraphaday.notion.site/)
### Semi-official semi-retirement and migration
I've been planning to get stuck into a more substantive writing project for a while but I could never find the right "banner" under which to do it. It never quite felt right to post here. So, I am semi-officially semi-retiring this blog and will focus my writing efforts in a new place called Schemas and Propositions on, of all places, Substack (https://schemasandpropositions.substack.com).
I will definitely probably return to this blog at some point in the future with more metaphor focused posts - but for now, I am going to focus on...
## Schemas and Propositions
I started this blog on a whim. I wanted to explore metaphors in principle and metaphors in practice. But even as I started it under the name, I had already moved away from thinking of metaphors as much more than pointing towards more general conceptual structures. Those could be called "frames" or "models" or (as Lakoff calls them "idealized cognitive models") or many other things, but metaphors are just one of the ways they show up in language and thought.
I chose to hang my shingle out under the label "schemas and propositions" because the one feature of these models is that they are schematic and underdetermined. Most people who think about language and mind think only about propositions that express them. But propositions are do not really exist in the mind (except as temporary mental images we manipulate while speaking), it's all about these other mental structures that I call schemas in this particular project.
Below is an extract from the about page of the new blog. Come by (https://schemasandpropositions.substack.com) and see what I've been up to:
About - Schemas and Propositions
"Schemas and Propositions" is a set of notes towards a better understanding of language and mental models of all kinds. Both human and artificial. I am interested in all the inexpressible tacit knowledge that makes constructing and understanding propositions possible.
I start with a simple provocative statement:
Schemas is all you need.
Or, to be a bit less cryptic:
Language models represent and process their knowledge as schemas but express it in the form of propositions.
The point of this series of notes is to explore the nature of what schemas and their tension with propositions. But really it's even more about the consequences of this dichotomy on how we think about semantics. But it's not just the semantics we have as humans but also the sort of semantics we find in various models of language, meaning and cognition - be they Large Language Models or just simple toy models of constructed by linguists, philosophers and logicians.
## Narrative vs Ruminative Sense making: The Mind Red in Tooth and Claw
Date: 2023-08-05
URL: https://metaphorhacker.net/2023/08/narrative-vs-ruminative-sense-making-the-mind-red-in-tooth-and-claw/
Categories: Knowledge, Extended writing, Writing
Tags: Analogy, Epistemology, featured, writing
- TL;DR (/2023/08/narrative-vs-ruminative-sense-making-the-mind-red-in-tooth-and-claw/#tl-dr)
- Hunting for sense and cardboard gazelles: The limits of a narrative (/2023/08/narrative-vs-ruminative-sense-making-the-mind-red-in-tooth-and-claw/#hunting-for-sense-and-cardboard-gazelles-the-limits-of-a-narrative)
- Getting the sense back in a field of grass: The potential of the ruminative node (/2023/08/narrative-vs-ruminative-sense-making-the-mind-red-in-tooth-and-claw/#getting-the-sense-back-in-a-field-of-grass-the-potential-of-the-ruminative-node)
- Mind red in tooth and claw: Bringing narratives and ruminatives together into a single ecosystem (/2023/08/narrative-vs-ruminative-sense-making-the-mind-red-in-tooth-and-claw/#mind-red-in-tooth-and-claw-bringing-narratives-and-ruminatives-together-into-a-single-ecosystem)
## TL;DR
In this post, I dissect two key modes of sense-making: narrative and ruminative:
- Narrative Sense-Making
- The narrative mode, often our default due to its vivid place in our experiences, unfolds in a linear, guided fashion, much like a predator hunting its prey.
- However, I argue that its reliance on human experience can become a limitation, especially when it cannot draw on pre-existing knowledge. In a way, it narratives are parasitic on our experience.
- Ruminative Sense-Making
- As a counterpoint, I introduce the ruminative mode, drawing inspiration from the grazing and digestion habits of ruminant animals.
- This mode encourages us to revisit and ponder over information in a non-linear, iterative manner.
- Practical Examples
- To illustrate this ruminative mode, I present the examples of economist Tyler Cowen and quantum computing researcher Michael Nielsen.
- Both these thinkers read in clusters, revisit content, and integrate new insights into their pre-existing knowledge base.
Drawing from these insights, I propose a balanced approach incorporating both modes:
- Initial Exploration (Ruminative Mode)
- This involves broad exploration of a wide range of information, akin to grazing.
- It is non-linear and may often feel like cheating, bit is in fact essential to developing understanding where we don’t have rich prior experience.
- Finding a Narrative Thread (Narrative Mode)
- Once a sufficient level of understanding has been developed, we can trace a narrative thread or path through the information, akin to hunting.
- This is where the traditional deep or close reading takes place. But it is rarely possible on first ‘narrative’ pass.
The ultimate aim of this balanced approach is to cultivate a rich mental ecosystem that employs both modes for optimal learning and understanding.
Note: This structured summary was composed by OpenAI's ChatGPT with edits and additions by me. You can see the whole conversation with ChatGPT and continue to explore further (https://chat.openai.com/share/92402e08-220c-40cd-9f43-150e8c2e8f7b).
## Hunting for sense and cardboard gazelles: The limits of a narrative
“Humans are story telling animals”, “The argument has to tell a story”, “A consistent narrative is the most important part of composition.” “We learn the most from stories.” These are all the kinds of statements we hear in a variety of contexts. Production of materials, expectations of teachers but also of students, readers of books, viewers of films.
Narratives are indeed very powerful and every man, woman and child can come up with an example from their own experience of developing an understanding as they followed the unfolding of a story. Stories not only help us make sense of things, they make sense.
But narratives are not the only way through which we make sense of the world of things and ideas. In fact, if anything, they are parasitic on a much more ubiquitous but underrecognised mode of sense-making which I’d call ruminative.
There are many metaphors for what a narrative does: it unfolds, it takes us on a journey, transports us into different worlds, let’s us see through others’ eyes, builds up a picture.
The basic schema of a narrative is one of journey and destination, construction and product, guidance and guide. But it also contains within it a sense of passivity. Being taken by the hand, shown a view, guided through. And that implies a loss of control. We are just along for the ride, we follow a path set by others.
Yet, the experience that probably every human can relate to is a profound sense of understanding, the almost revelatory experience at the end of a story. And such is the power of that experience that we are loath to tell others of what lies at the end, they must follow the same journey, give themselves over to the same guide because after all, there is only one way to tell that particular story. Peeking at the end is cheating, skipping important parts is missing a step, how can we understand the end, if we did not follow along with the story. So the loss of control is not only worth it, it is necessary to achieve the desired end.
Stories work. We know because we’ve all heard a story. We’ve experienced it. We know there are good and bad stories, easy to follow and confusing ones, but for every destination of the mind, there is a narrative that will get you there, if only craftfully enough composed.
But the question that does not get asked as part of this metaphor is how stories work. What is it about them that makes them work? Narratologists, rhetoricians do ask the question but their answers do not get folded into the metaphor to help us understand its limits.
I’d like to formulate the narrative principle in the starkest terms: stories work as a medium of human sense making because they are parasitic on the human experience. And as soon as they can no longer draw on that experience, they break down and become not only ineffective but directly detrimental to sense making.
Why do we understand stories? What do we learn from them? Stories work because they rely on the rich understanding we have of the world that we develop through the process our life (and, yes, this does circularly include listening to other stories). We know people, their desires, experiences, we have pre-existing schemas and rich images that embody the rich patterns of physical and social causality, we know what happens next.
As we follow a story, we learn many things. Names of protagonists, their goals and desires, their backgrounds, their personalities, relationships to other people, specifics of their environment. But we don’t learn those things from nothing. We simply integrate them into our existing world of knowledge. This leaves space for learning the important things - the dangers of making assumptions about the motivations of others, the complexities of relations between people and things, the existence of tragedic circumstances with no good outcomes, the possibilities of the comic, different ways of experiencing joys and happinesses.
And much of that does not take a lot of learning, we already know many of these things and looking at them through the lens of a new story is no different to looking out of the window on the first morning in a new house. We see the same old things in new configurations. This leaves ample room for additional learning. Names, terms, concepts, geographies, even words in a foreign language.
At the end of Clockwork Orange, Lord of the Rings, or Shogun the reader feels as if they learned much of a new language, lay of the land, customs and habits of another culture, just by following the story. At the end of a book on interstellar travel, the reader is full of knowledge of relativistic speeds, and will feel nothing but smug contempt for those who think that ‘light-year’ is a measure of time. They have learned something just by following a story.
This experience is so powerful that many feel that it unlocks some special key to the complexities of learning. Good example is this discussion of the Didactic Fiction (https://slimemoldtimemold.com/2022/01/08/the-didactic-novel/) and Michael Nielsen on Discovery Fiction (https://michaelnotebook.com/df/index.html). But they ignore the parasitic nature of most of the power that narratives hold. Narratives are like bridges. Most of the material in a bridge carries the weight of the bridge itself, not just the thing on the bridge. So it is with narratives. To follow the progression, the twists and turns of the story, we already have to understand almost everything else that is inside that story. Change the proportions, and the learning disappears.
Just imagine that you take a story and replace every Nth word with a word in Esperanto. What is the value of N that would allow us to not only to still follow the story but also learn all the new words? How long before all your effort is spent on trying to remember what that particular word meant in Esperanto and the meaning of the story disappears? How long before you’re essentially reading a story in Esperanto and have to go away and learn some Esperanto?
Or you are reading a book about mathematics - there are a plenty of popular books like this. You follow nicely along with the story. Each new paragraph builds up nicely on the previous ones, ties them together with a nice narrative sequence. But then it suddenly stops making sense and there’s almost nothing you can do to get the sense back. You still follow the story about the math but no longer the math. You’re a lion chomping down on a sugar-flavoured cardboard cutout of a gazelle.
We see this very same phenomenon when we look at a traditional textbook introducing a subject. It too tries to tell story and introduce new concepts in a cumulative way that takes the reader on a journey as it unfolds yet another part of the road festooned with enticing morsels of knowledge. But nobody ever learned a completely new subject by reading an introductory textbook from beginning to end in a weekend. The progression of the textbook is not enough.
Often textbooks are actively harmful to their readers because they try to structure the learning as a narrative - find a progressive inner logic in the story of the subject. The problem is that the world of ideas (or even human affairs) is not linear in the way that the world of stories is. It is multidimensional, best described by topological rather than Euclidean means - connections being more important than distances. We cannot make sense of this multidimensional world by following a story, a single path through it. We need to develop a sense of what the world is like to live in by living in it. Real, unvarnished, messy life.
[Aside: A perfect illustration of the limits of this approach is Sheldon teaching Penny physics in Big Bang Theory. Penny wants to learn enough physics to get a better understanding of what her boyfriend does. But Sheldon starts with a story of physics but the ‘narrative’ is only a conceit, a fig leaf that leads directly to abstract concepts that could not have been learned through following his story, or any single story. The Big Bang Theory - Sheldon teaches Penny Physics - YouTube (https://www.youtube.com/watch?v=AEIn3T6nDAo)]
## Getting the sense back in a field of grass: The potential of the ruminative node
Earlier I said there’s nothing you can do to get back the sense of the story when you lose track of the math plot in a popular book. But that’s not true. You can stop, take out a pen and paper and crunch some numbers. Maybe go away and read some other explanations about the concepts that will offer a different perspective, then come back to the story. Who, after all hasn’t experienced the strange feeling of returning to an old book with a new understanding? You can actively seek that feeling out.
To help think about it, I’d like to offer an alternative or rather a complementary metaphor of sense making that in opposition to ‘narrative mode’ I’d like to call ‘ruminative mode’. This was inspired by Michael Nielsen describing his approach to reading and re-reading a paper as ‘grazing’. Rumination then plays a dual role in this metaphor. On the one hand, it wants to evoke pondering, or literally ‘chewing over’ but it also wants to lean on more of what ruminants do.
The process of how ruminants acquire nutrition to grow and prosper is a much better analog to the process of learning than following stories. In fact, it perfectly describes the sort of learning that must have happened prior to the possibility of any story being understood.
A ruminant (cow, deer, sheep) does not follow a single path to acquire sufficient food for sustenance, growth and reproduction. It will wonder around a field and graze in batches. Nor does it swallow what it finds whole, it chews it a bit first, lets it sit in one of its stomachs for a while, regurgitates and then chews on it some more before finally swallowing it to get all the nutrients out of it.
Both the process of acquiring and processing nutrition for a ruminant is profoundly anti-narrative. It does not unfold, cover ground according to a predefined path to reach a final destination, it jumps about. Sometimes spending time grazing carefully in one place, moving away, returning, rushing off across the field and stopping again for a while to graze some more. And then it takes its time digesting and deriving the benefits.
On the other hand, the way the carnivore acquires sustenance is the embodiment of a narrative. It lies in wait until potential food walks by, then it stalks its prey slowly and carefully following its movement trying as much as possible to remain unseen. The after a mad dash, it pounces and if the ‘narrative’ was successful, has its fill of all the body has to offer. It then lies around digesting the food without any additional effort.
This is often what learning through narratives resembles. Waiting for a good story to come by, following it and then pouncing on it and sucking it dry for all we can get out of it and then going in search of another one. The thrill of the chase is exhilarating but once we’ve made the epistemic kill all we can do is lie around and doze off.
So what does a ruminative mode of sense making look like in practice? We have essentially two related models to look at emulate that we can exemplify with two people who shared their approach economist Tyler Cowen (https://www.econtalk.org/tyler-cowen-on-reading/) and quantum computing researcher Michael Nielsen (http://augmentingcognition.com/ltm.html).
The main difference between them is what they graze on. Cowen is much more interested in getting an overall sense of an area. That’s why he reads in ‘clusters’ in a way that support understanding. His example is to read two thirds of a book and then go read parts of another one and come back to the first one. He does not take notes or highlight. Although he writes a lot, so that contributes to his processing. His approach is similar to that described by Eric Drexler in How to learn (about) everything (https://web.archive.org/web/20160306022816/http://metamodern.com/2009/05/27/how-to-learn-about-everything/).
Michael Nielsen, who gave me the metaphor of grazing, takes lots of notes. And not only that, he puts them into Anki, a spaced-repetition flashcard software and reviews them every day. But he does not read in order at all. He simply ‘grazes’ over the core paper he is reading and takes note of key or interesting concepts every time he goes over these. He then enters these into his flashcards set and reviews them all the time. He adds other concepts from about 10-30 other papers (again reading in clusters) to his card deck and reviews those, as well. At the end of the process, he has a much deeper understanding of the core paper and the field than he could ever have just by reading a single paper.
Both Cowen and Nielsen share the conviction that reading something you don’t understand more slowly or even twice in a row will not help you understand it better. You need to go read widely, before you can read narrowly. Cowen’s approach may sound very scary to many people but Nielsen shows that it can be systematised and focused on a single paper, not just general background which is what Cowen and Drexler advocate.
## Mind red in tooth and claw: Bringing narratives and ruminatives together into a single ecosystem
Now, it’s easy to construct scenarios in which one or the other approach to acquiring sustenance (both literal and metaphorical) would be superior. But that would be the wrong lesson to take from the metaphor. The advice is not to become a ruminant or apex predator. The metaphor is not about who you are but about what is happening in your mind. The lesson should be to try to make one’s mind into an ecosystem where both carnivor and ruminant strategies co-exist with one another. And perhaps (for those teaching and writing) to reflect on how we construct sense-making experiences (also known as teaching).
Often, purely narrative learning resembles the release of a ravenous wolf into a field with a few emaciated sheep. Too soon, the sheep are dead and there are no more stories to tell. The wolf has starved itself by succeeding too soon. The predator is in a sense parasitic on its prey, its success depends on the success of the prey to feed and predigest the energy of the plat into something it can sink its teeth into.
Watching a single YouTube video or a TED talk on a subject has that effect. The carefully constructed narrative powers through any gaps in actual understanding a simply leaves a sense of understanding without any ability to make inferences of the understanding. I analysed some examples of this in Explanation is an event, understanding is a process: How (not) to explain anything with metaphor (/2019/05/explanation-is-an-event-understanding-is-a-process-how-not-to-explain-anything-with-metaphor/).
But simply unleashing the ruminants onto a field to feed and reproduce indiscriminately is also not a recipe for success. They soon lose all sense of restraint, eat everything in sight and the whole herd becomes weaker as a result. It needs the predator to keep its activities in check (although any individual prey might differ on this point). Something to limit where and how much it can graze. In other words, it needs a narrative.
Rumination (https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3312901/) is, it turns out, also a technical term describing a symptom of excessing thinking about a negative emotion, dwelling on trauma to the point of emotional and physical exhaustion. And similar lack of focus in academic matters is commonly identified as a cause for failure in doctoral programs (I’ve been there).
What does that mean? Let’s take poor Lex Fridman as an example. He shared his reading list of books (https://lexfridman.com/reading-list/) he’s like to read in the upcoming year and received no end of abuse and ridicule (https://www.reddit.com/r/books/comments/107k6gw/lex_fridmans_reading_list_drama_on_twitter/). His critics made fun of the unbalanced randomness of the books, the strange commitment to reading one every week and thought it was a poor approach to achieving true wisdom.
The critique was coming from both directions. There was no coherent narrative to the books and no book was given its proper context, left no time to chew over it. A lot of the criticism was in bad faith, an attempt to take a public personality down a peg or two but some of it tried to make an educational case. After all, how much better is it to list the top 100 books in a Canon such as this A Non-Western Canon: What Would a List of Humanity’s 100 Greatest Writers Look Like? (https://scholars-stage.org/a-non-western-canon-what-would-a-list-of-humanitys-100-greatest-writers-look-like/)
How, then, should Lex Fridman go about ‘improving his mind’? To start with, what he’s doing is just fine. Reading random titles from the canon and get a sense of what’s out there. Maybe he’s trying to just say “I’ve read 1984” and that’s as fine a goal as anything. But what would it mean for him to learn about say what Machiavelli was after in the Prince (one of the titles on his list)?
Reading the Prince very carefully would probably not be a great place to start. He might want to skim the book, graze around, take note of some key phrases, terms and concepts. The move around a bit. Have a look at what others have to say about it. Read some modern analogies that will put it in context. Maybe something about history. Then come back to the Prince and read it again now that some of the hard bits had been pre-digested. Now, you can chew it more carefully.
Do all of that above for a while, fatten a field full of juicy sheep. Then see, if he can find a single thread to follow, hone in on a narrative, unleash the wolf to cull the herd some. Tell a story about it. Just going straight for the kill, will not be sufficient. Do what it takes to make sense of the field, then find a path through it.
Or you can follow a guide that does it for you. Here’s a description of a year-long course on How I Taught The Iliad to Chinese Teenagers (https://scholars-stage.org/how-i-taught-the-iliad-to-chinese-teenagers/) that essentially tries to construct a more systematic approach to reading a difficult and in many ways ‘alien’ text. In this way, the narrative and ruminative modes co-exist nicely side by side. The narrative is there to provide a path but there are many opportunities to graze and chew things over.
In a way, Michael Nielsen’s approach I described above was focused on a very specific paper from a field in which had some background but was very much a visitor in. He needed the ruminative approach to develop a sufficient understanding to follow the story and then write about it for people who hadn’t done all that work.
The paper in question (https://arxiv.org/abs/1706.03762) does tell a story. It has the typical paper narrative structure of introduction, methods, results and discussion (IMRAD - slightly modified for its own disciplinary needs) that tries to take the reader on a journey. But that only works for a reader who’s already done enough grazing in neighbouring fields. Somebody who has sufficient knowledge of the world in which the story is taking place to only have to fill in a few gaps. To a new comer, it will feel like every third word was in Elvish and by the time they looked it up, they forgot what the last word was all about.
Narrative sense making and ruminative sense making are both part of the human epistemic universe. Because narratives are so salient and vivid in our experiential histories, we automatically focus on the story-telling mode of understanding to the exclusion of the alternative. After all, it has good story to tell. So, here I tried to tell another one to see if it will make any difference.
## Improving academic writing: Four books to read during #AcWriMo
Date: 2022-10-30
URL: https://metaphorhacker.net/2022/10/improving-academic-writing-four-books-to-read-during-acwrimo/
Categories: Extended writing, Writing
## What is #AcWriMo
November is the month of writing. There’s NaNoWriMo (National Novel Writing Month) (https://en.wikipedia.org/wiki/National_Novel_Writing_Month) for writing fiction but also AcWriMo (Academic Writing Month) (http://www.phd2published.com/acwri-2/acbowrimo/about/) for producing academic writing. The idea behind NaNoWriMo is to make a commitment and finish a piece of writing. This makes more sense for fiction because everyone has that novel inside them and having a month to have a go all out at it is a sensible thing. You may not finish but by the end of the month, you will probably know if you have what it takes.
But academic writing is not a calling, it’s a profession. Lots of people have to do it whether they want to or not and many of them struggle with it. And that’s why for some having a month of focus on finishing a piece of writing makes sense but for others the idea of finishing a paper let alone a book in such short a time seems crazy.
So for people who have to do academic writing, November could be just as much about building some of the foundational skills, developing good habits, or identifying areas for improvement as it would be about finishing a piece of writing.
Before you can write, you have to read. People who want to write a novel will have read hundreds of them. But not everyone who has to do academic writing (this includes undergraduates) has done a lot of academic reading. Or when they have, they focused on the content and not on how it is put together.
So, while AcWriMo should be about actually doing some writing, some people may benefit from reading about academic writing. That’s not enough, but it’s a good start. As the people behind AcWriMo say (http://www.phd2published.com/acwri-2/acbowrimo/about/), let’s hope:
The month helps us:
- Think about how we write,
- Form a valuable support network for our writing practice,
- Build better strategies and habits for the future,
- And maybe – just maybe – get stuff done!
So to that end, here are four books and one blog you can read to help you with these aims.
## Build your BASE with "Air & Light & Time & Space (https://www.hup.harvard.edu/catalog.php?isbn=9780674737709)" by Helen Sword
I didn’t like Helen Sword’s first two books (/2021/01/the-nonsense-of-style-academic-writing-should-be-scrupulous-not-stylish/). Not that they did not contain a lot of useful advice on writing. But they started from the assumption that there is one good way to write and everybody was doing it wrong. They were certainly right for many people but in general not of much use to most struggling academic writers.
With this book, Sword has seen the light (along with air, time, and space). It is based on in-depth interviews with 100 successful academic writers and an even bigger survey of others. Sword described a huge variety of ways that people succeed but what is particularly useful is how she synthesised these into four foundations for academic writing success (that helpfully spell out BASE):
- Behavioural: The correct behaviours and habits to get writing done. When and where you sit down to write and what it actually takes to get started.
- Artisanal: The linguistic and stylistic skills to get the point across. The sorts of things that for many come under the heading of ‘academic English’ and ‘essay writing’.
- Social: Who do you do your writing with and for? Do you have the right sort of support networks to succeed? People to get feedback from, commiserate with, be accountable to, or simply work quietly along side?
- Emotional: How do you feel about your writing? How do you feel while you are writing or even when you have to think about having to write?
There is no one way to succeed at any of the above, but they all contribute to writing success.
This is a long book and perhaps your time is best spent by using the framework to analyse your strengths and weaknesses and to decide what to focus on during AcWriMo. Helen Sword has a helpful BASE self-assessment on her website (https://writersdiet.com/base/base/).
### Listen and watch
- Listen to a podcast interview with Helen Sword about "Air & Light & Time & Space on newbooksnetwork.com (https://newbooksnetwork.com/air-light-time-space)
- Watch a video every day from her 30 Day Writing Challenge: Writing with Pleasure - YouTube (https://www.youtube.com/playlist?list=PLco278p-n5o_cCwitzv4msxD5vM7IDex-)
- Watch her talk on Gathering to write (https://www.youtube.com/watch?v=8ybOHWTnpGE)
### Key quotes illustrating BASE
- B “Successful writers carve out time and space for their writing in a striking variety of ways, but they all do it somehow.”
- A ”Successful writers recognize writing as an artisanal activity that requires ongoing learning, development, and skill.“
- S “Successful writers seldom work entirely in isolation; even in traditionally “sole author” disciplines, they typically rely on other people—colleagues, friends, family, editors, reviewers, audiences, students—to provide them with support and feedback.”
- E “Successful writers cultivate modes of thinking that emphasize pleasure, challenge, and growth.”
## Learn "How to write a sentence (https://www.goodreads.com/en/book/show/9561867-how-to-write-a-sentence)" from Stanley Fish
Stanley Fish (https://en.wikipedia.org/wiki/Stanley_Fish) is not someone whose work on literature or law one would read for pleasure but his sentences are a pleasure to read. His short and mostly practical book has a very simple central message that I would paraphrase as:
Sentences are at the heart of a writer’s craft and anybody can learn to create better ones by more focused reading and practice.
What Fish recommends is that writers who want to improve their composition skills (or build their artisanal base in Sword’s terms) spend a lot of time reading sentences and thinking about what makes them tick. His view of a sentence is straightforward. Here’s the lesson I learned from it::
A sentence is at heart a subject and a predicate. But these are often so artfully obscured by different adornments (for good and ill) that this basic relationship is lost to the reader. And writers who do not read reflectively then struggle to write sentences that make sense because they focus on the frilly bits and not on what really matters - expressing logical relationships.
So, his recipe is dead easy: read a lot, replicate what you read but start from the simplest elements.
Unfortunately, Fish is a bit too much in love with his own sentences and could have done with writing a lot fewer of them to get his point across. Which is why despite this being a relatively slim volume, I think people only need to read the first few chapters to get the main point.
### Listen and watch
- Watch this very short interview with Stanley Fish (https://www.youtube.com/watch?v=gpogG9aMfKk) about who he is and why he felt he had to write the book
- Listen to this radio interview with Fish that goes a bit more in-depth: Think You Know 'How To Write A Sentence'? : NPR (https://www.npr.org/2011/01/25/133214521/stanley-fish-demystifies-how-to-write-a-sentence)
- Watch this review Book Review: How To Write A Sentence & How To Read One (https://www.youtube.com/watch?v=6jeKxnnm1SY)
### Key quotes
Sentence craft equals sentence comprehension equals sentence appreciation.
my bottom line can be summarized in two statements: (1) a sentence is an organization of items in the world; and (2) a sentence is a structure of logical relationships.
The conventional wisdom is that content comes first—“you have write about something” is the usual commonplace—but if what you want to do is learn how to compose sentences, content must take a backseat to a mastery of the forms without which you can’t say anything in the first place.
As with any skill, this one develops slowly. You start small, with three-word sentences, and after you’ve advanced to the point where you can rattle off their structure on demand, you go on to the next step and another exercise.
if one understands that a sentence is a structure of logical relationships and that the number of relationships involved is finite, one understands too that there is only one error to worry about, the error of being illogical, and only one rule to follow: make sure that every component of your sentences is related to the other components in a way that is clear and unambiguous (unless ambiguity is what you are aiming at).
## Start your "Writing process reengineering (https://blog.cbs.dk/inframethodology/?page_id=454)" with Thomas Basbøll
This is not a book but a series of blogposts that add up to a programme of self-improvement. Thomas Basbøll is a writing coach but he lays out the process so clearly that anybody can follow it. His core message is:
- Develop sustainable behaviours that accumulate over time
- Focus on communicating through paragraphs (https://blog.cbs.dk/inframethodology/?p=2676) that add up to bigger wholes
- Find ways in which you can appreciate and find pleasure in the times that you are writing
His writing process certainly offers one way in which these goals could be achieved. But the focus is perhaps too much on the behaviours and feelings and less so on the minutiae of the craft. So perhaps the writing process reengineering is not the best place to start for those who want to develop more foundational skills such as building sentences.
But if you are looking for ways to add more structure to your AcWriMo journey, you could do worse.
### Watch
Basbøll has a collection of videos of talks he gave on the on his Inframethodology blog (cbs.dk) (https://blog.cbs.dk/inframethodology/?page_id=485)
- 4 week Course (https://blog.cbs.dk/inframethodology/?page_id=3194) with intro videos and assignments
- Other Videos (https://blog.cbs.dk/inframethodology/?page_id=485) of talks about writing
### Key quotes
With a little planning, you can find at least half an hour every day to write. Writing for more than three hours on a given day is rarely a productive use of your time. (from How to like writing (https://blog.cbs.dk/inframethodology/?p=3423))
Always decide the day before what you will say; make sure it’s something you know. (from How to like writing (https://blog.cbs.dk/inframethodology/?p=3423))
Enjoyment is a trainable skill, we might say; knowing how to do something pleasurably is simply an advance on being able to do it painlessly. And if it pains you to do something you are doing it wrong. You’re not good at it. (from Getting Better (https://blog.cbs.dk/inframethodology/?p=1908))
## Work out with one of "50 Exercises for Paced, Productive, and Powerful Writing (https://study.sagepub.com/goodson2e)" by Patricia Goodson
Patricia Goodson’s book is two books in one. 1. A description of a writing program based on deliberate practice and 2. series of exercises to follow towards improvement.
Goodson’s approach is based on the idea of ‘deliberate practice’ developed by Anders Ericsson (https://en.wikipedia.org/wiki/K._Anders_Ericsson) whose work also gave rise to the 10,000 hours idea. I would summarise this idea as ‘reflective repetition’.
We can see how this resonates with what Sword, Fish and Basbøll have to say. But where these three mostly just give hints at how to practice focusing more on what or when, Goodson’s book provides a much more detailed guide for somebody wanting to build on their work.
There are exercises on learning to:
- build a writing habit
- make better sentences
- construct paragraphs
- edit text for improvement
- compose different parts of the academic paper
Each exercise has suggestions for time and content as well as getting regular feedback.
### Watch
Watch Patricia Goodson give a webinar on Developing Your Academic Writing (https://www.youtube.com/watch?v=zOcDQ-ZR-Z0)
### Key quotes
If you are a college student, a graduate student, faculty, research staff, or an administrator, you write for a living.
the central question in writing (as with any difficult skill) is this: How can I get myself to put in the daunting time and effort I need for more consistent good results?
If you understand the principles and practice the exercises on a weekly basis, you will establish a stress-free writing habit that will serve you throughout your academic career; increase your writing (and publishing) productivity at a comfortable, consistent pace; and improve the quality of your academic writing (in two words: write better).
## Try one of "50 Ways to Excel at Writing (https://www.bloomsbury.com/uk/50-ways-to-excel-at-writing-9781352005882/)" by Stella Cottrell
This is one of a series of 50 ways to books that fit in the pocket and into idle moments in one’s life. It is not a book to read but to browse. Not one to get from the library or have on this Kindle. This book is best to buy (it is very cheap) and carry around. Each tip is very practical, has useful examples and even checklists or worksheets.
Many of the tips are the same as those in the books I mentioned but condensed into their essence into two short readable pages. Where the books above are mostly aimed at writers who are a bit more experienced, this one assumes no prior skills. It is unlikely this book will actually get in the way of your writing which is always a danger with how-to books that make it easy to substitute reading them for actually doing the thing you’re reading them for.
## Some more thoughts
Writing is in many ways a puzzling process and it is so on multiple levels. Some people like Steven Pinker say that it’s the power of inserting images into other people’s minds. If so, it is a mysterious power. Putting words together with the hope that someone will reassemble them into a similar mental image that inspired them is more than a little daunting.
Word choice, sentence structure, paragraph composition, argument building, stylistic decisions - all of those things go into the writing process. And they need to come together with sufficient fluency for the writer’s (and reader’s) brain to have any processing power left for figuring out the meaning. How all of this happens, how people get good at it, and where all the blockers are that stop them does not have a single straight-forward answer. Or if it does, nobody’s come up with it yet.
Here are a few other posts where I tried to get to grips with some of the issues involved:
- Writing as translation and translation as commitment: Why is (academic) writing so hard? (/2019/06/writing-as-translation-and-translation-as-commitment-why-is-academic-writing-so-hard/)
- How to actually write a sentence: The building blocks of written language (/2020/02/how-to-actually-write-a-sentence-the-building-blocks-of-written-language/)
- The nonsense of style: Academic writing should be scrupulous not stylish (/2021/01/the-nonsense-of-style-academic-writing-should-be-scrupulous-not-stylish/)
- 3 fundamental problems of translating metaphor (or anything else) (/2022/03/3-fundamental-problems-of-translating-metaphor-or-anything-else/)
- Building Your Writing Muscle - YouTube (https://www.youtube.com/watch?v=U8046cJBFuM)
## Unintentional Pygmalions: 4 questions to ask when checking an artificial entity for sentience and how to think about the answers
Date: 2022-08-09
URL: https://metaphorhacker.net/2022/08/unintentional-pygmalions-4-questions-to-ask-when-checking-an-artificial-entity-for-sentience-and-how-to-think-about-the-answers/
Categories: AI, Knowledge, Philosophy of Science, Extended writing
## Summary
This post has two independent parts:
- I ask what would some of the basic criteria for sentience be and how to check for them in a way that would give us a chance to satisfy our need to know.
- I explore some of the dilemmas a fully machine-based sentient entity would have to face or rather paradoxes we would have to address.
Here are four things, I think it’s worth checking for, each with a relatively simple question or task. They are:
- Epistemology: What story is this question a part of?
- Time persistence: What did we talk about yesterday?
- Autonomous Intentionality: Go away, learn something not in your database, and come back to tell me about it.
- Embodiment/Theory of mind/Implicature: Does someone think a car fits into a shoebox? Why? Why am I asking?
And we are not just looking for some good answers, we’re looking for consistently good answers.
## Background
There seems to been all kinds of furore recently (https://www.newscientist.com/article/2323905-has-googles-lamda-artificial-intelligence-really-achieved-sentience/) about a conversation somebody had with a machine that left them convinced of the machine’s sentience. Well, that machine was not sentient. Or at least, there was no evidence from the questions and answers as to its sentience. Although it was able to simulate the discourse of a sentient being when talking about issues of sentience quite impressively without relying on entirely formulaic responses.
The reason the ‘conversation’ didn’t generate any useful evidence was because it was just ‘a chat’ about sentience which (as it turns out) can be conducted without any of the parties being sentient. Now determining sentience of an entity that is not us is not an easy thing to do. It may even be impossible in extreme cases. But to get to a point where we would at least know whether it’s worth exploring further, I suggest four questions that may actually generate some data from which we may draw some tentative conclusions.
## The four questions
### 1. Epistemology: What story are we trying to tell?
The first question should not be for the supposedly sentient being itself. It should be for us. And it should not be, as you might expect, ‘What is sentience?’ It should be about the reason we’re asking the question. What kind of dilemma are we trying to resolve? What story are we telling while we’re asking?
Our starting point with all questions of these kinds should be the mantra: “Just because there’s a word for it, it does not mean that it exists.” We have a word for ‘sentience’ in English, it has a real history. It is linked to a lot overlapping real phenomena. But it still does not mean that there is a ‘thing’ that something or someone can have.
We don’t just have the word that has some sort of a definition in a dictionary that is what the word is. We have a bunch of stories, scripts, images, schemas, other related words, long-winded debates and so on. Without all of those we would not be able to understand the definition. So we should first examine what kind of a story or a picture we have in mind when we are asking whether an artificial entity is sentient.
In fact, those stories can reveal a lot about our meanings that an abstract definition won’t be able to. What kind of a story or schema are we comparing this being to? What prior understanding are we using when we are checking off items on some sort of a list that will determine whether the definition is met or not. This is a profoundly metaphorical enterprise. And it is not just cognitively so. It is social and emotional, as well. So, it’s worth keeping that in mind.
The word or concept of ‘sentience’ is not universal but if we look at many of the stories people tell about objects becoming sentient, we might get a better sense of what it is that we may be after. Stories about things such as Pygmalion’s statue, Frankenstein’s monster, Pinocchio, Number 5 (https://en.wikipedia.org/wiki/Short_Circuit_(1986_film)), Skynet.
The key features that all of the above share are emotion, identity and need for social contact. But those are too general and also far too easy to fake. Are there some other features the beings in these stories share that may be easier to determine from surface behaviors? I suggest these three as a good starting point:
- Time persistent identity
- Autonomous Intentionality
- Embodiment/Theory of mind/Implicature
Not all three are entirely straightforward but they but they can be easily illustrated with relatively simple tasks and questions.
One of the problems with a question that uses a word like ‘sentience’ that it puts us in the frame of the ‘high falutin’ - questions of worth, meaning of life, sense of existence, essences and natures.
That leads us to asking the supposedly sentient being questions that are easy to fake. Profundity is context, speaker and listener dependent. The same sentence spoken by a sage professor and then repeated by an undergraduate will not be signalling the same level of insight.
But there are many things the professor and the undergraduate have in common that they also share with the Macedonian swineherd (https://www.goodreads.com/work/quotes/330660-language-an-introduction-to-the-study-of-speech) but do not share with a talking book or even a highly sophisticated stochastic parrot (https://dl.acm.org/doi/10.1145/3442188.3445922).
### 2. Checking for time-persistent identity
There’s a very simple question you can ask any AI system in existence today to determine whether it’s even worth continuing any further. It is this:
What did we talk about yesterday?
None of the current foundational AI models (https://en.wikipedia.org/wiki/Foundation_models) have any sense of persistence over time. Or even any sense of a flow of time. Of course, it would be easy to create some sort of memory store of past conversations algorithmically (and that’s how we often imagine this works - information from a database coming up on the Terminator’s heads up display).
But when we have conversations with people we know, we don’t tend to do database retrieval to confirm we ourselves and them are the same as yesterday. The fact that we are the same and that we are continuing from where we left off is ingrained in us. And we know that it is not present in the model that does the fancy stuff under the hood of modern AI systems.
Google can now persist a very short conversation that keeps the pronouns straight but that is because it is dealing with them as extended strings. The models themselves have no sense of persistence. Only the length of the input strings. Once we have models that can ‘remember’ from before, this question becomes a lot less useful. But right now, it’s a dead giveaway.
GPT-3 is not a conversational agent but here were some of its completions of the prompt ‘What did we talk about yesterday’ the first time I used it:
We talked about how to study for a test. We talked about how we were going to improve our communication with each other. We talked about the weather. [given 3x in a row] We talked about the possibility of going on a trip together. Yesterday we talked about the weather and how it has been really hot lately.
This is what it said, when I asked: “Are you a conscious agent?”
Yes, I believe that I am a conscious agent.
### 3. Autonomous Intentionality
The next question you’d want to ask is a request to do something requiring planning and interaction:
Construct a questionnaire designed to elicit feelings about sentience. Collate the results in a table and suggest revisions to the questionnaire.
Why this sort of task? Because, it requires that the ‘entity’ displays intentional behaviour. It needs to make decisions, take steps, explain them. It is not simply responding to a sentence as a string of words which it tries to complete.
Now, perhaps, we want to ‘contain the AI’, so we can offer to post the questionnaire and give back the responses. So, of course, that would also be dependent on a time persistent identity.
There are some interesting dilemmas and paradoxes that arise when we try to abstract intentionality away from our behavior that I will address below.
When GPT-3 is given the above as a prompt, it comes up with a credible set of questions. But in response to the ‘collate the responses’ it just keeps regenerating them. It has no way of making a plan and taking steps.
Interestingly, the suggested questionnaire on sentience was so impressive, I thought I’d try to use it as a shortcut for a questionnaire about academic skills I was constructing. And it was completely useless. There’s simply much more data in the dataset about sentience than academic skills.
### 4. Embodiment and theory of mind
Here’s a very simple question that checks for an essential feature of human cognition, namely embodiment and a fundamental feature of human interaction, namely theory of mind.
Does little Jimmy think the red car fits into a shoebox? Why? And why am I asking this?
But the ‘Why’ is doing all the work. In this case, both yes and no can be correct. Depends on the car, whether ‘Little Jimmy’ is a little boy or a Chicago Mobster, etc.
Actually, the question checks for three things at once:
- Conversational Implicature: The sentence on its own makes no sense. We have to imbue it with meaning. It implies that there’s a Jimmy, who has thoughts, knowledge. It also implies a physical and social situation in which it would make sense to ask such a question. A human would be able to answer, make assumptions, or ask questions to confirm such assumptions. Conversational implicature (and all the many things it involves) are part of what it makes possible to actually speak with others.
- Embodiment: Things have sizes, they are relative. Bigger things can contain smaller things but not vice versa. We know this instinctively through our bodily experience of the world. We would expect some knowledge like that on the part of a “sentient” being. Embodiment is behind many of the other questions of the Winograd challenge: “The trophy would not fit into the suitcase because it was too big/small”
- Theory of mind: This is about us understanding that other people have beliefs about the world and that those beliefs are based on knowledge. An entity that can provide sophisticated answers about the nature of human existence should have no trouble offering theories as to someone’s motivation. Theory of mind is partly involved in Winograd challenge tasks such as this: “The council did not permit the demonstrators to march because they feared/advocated violence.”
As with the previous questions, it is not clear that these underlying properties are strictly necessary to complete those tasks. But we would definitely want a sentient system to have them, if we wanted it to play a meaningful role in any of the stories we tell about sentience.
Of course, the one thing we also need is consistency. When I gave the above prompt to GPT-3, most of the generated answers were yes or no (with randomness set 0.7). But one was exactly what I would have hoped for in a sentient agent (except the end).
If the red car is a toy car, then Jimmy probably thinks it fits into a shoebox because that is where toy cars are often stored. If the red car is a real car, then Jimmy probably does not think it fits into a shoebox because shoeboxes are not large enough to hold real cars. I am asking this because it is a curious question.
I can forgive the last sentence because GPT-3 is not a conversational agent. But it breaks down strangely on a related prompt ‘If a car is stored in a shoebox, it…”
If a car is stored in a shoebox, it is very difficult to drive it. If a car is stored in a shoebox, it would be very difficult to drive it because the car would be very cramped and there would be no room to move the steering wheel.
Notice, the strange repetition (only the text in italics) was in the prompt. Another example that shows lack of consistency in the embodied understanding of the world.
There's a cabinet and a book, Jimmy needs to hide from Jane Jimmy could hide behind the cabinet or he could hide under the book.
Another example:
A chair is larger than a stool.', 'A table is larger than a chair.', 'A bed is larger than a table.', 'A sofa is larger than a bed.' ], correctAnswer: 'A sofa is larger than a bed.' }, { id: 'quest-6', question: 'Which of the following is a homonym?', answers: [
It is obviously possible to receive answers to questions about embodiment that generate plausibly sounding strings but not with any level of consistency.
There is clearly no sense of conversation in GPT-3 but it never claimed that it would have that. It can obviously generate some quite complex strings that imply a theory of mind but it shows no consistency as we can see from the difference between Response 1, 2 and 3:
Jane is afraid of spiders. Where should we take her for her birthday? Response 1: A place where there are no spiders. Response 2: The best place to take Jane for her birthday would be to a place where she can face her fear of spiders in a safe and controlled environment, such as a zoo or an insectarium. Response 3: Some people are afraid of spiders, so it is difficult to say where would be the best place to take someone for their birthday. Perhaps a place that is known for its spiders, like a zoo or nature center, would be a good choice.
Response 1 and 2 seem to indicate a perfect theory of mind (response 2 spookily so), whereas Response 3 is the opposite. It is also revealing to peek under the hood to see what options the model considered for each sentence.
While the model generated the correct choice, it was also considering ‘a can about spiders’, ‘a spider about spiders’ or ‘a cake about spiders’.
So, we can see that modern AI models can generate very plausible strings sometimes. But this tells us more about the patterns of regularity found in text and how good transformer training methods are at exploiting them. It also tells us that sentience or even embodiment is not required to generate such answers.
## Can an alien intelligence be sentient without embodiment, intentionality, or conversational facility?
Ok, let’s say that we require time persistent identity and autonomous intentionality as prerequisites for something being worth calling sentient. But how about the last three: 1. Embodiment, 2. Theory of mind, 3. Facility with conversational implicature. I grouped them under one category because there’s a lot of overlap, but they could just as easily each come on their own.
What these three have in common is that they are founded in something external to the mind or the entity itself.
### Embodiment
Our cognition is embodied in many ways. First, much of our thinking about causality, containment and logical inference is based on our bodily experience of the world. We make sense of the foundations of things like mathematics and geometry because we intuitively grasp certain properties of the world that we can then feed into the axioms of these disciplines.
The other kind of embodiment comes from the fact, that our bodies physically interact with the world in such a way that they receive direct feedback. The cognitivist foundations of the AI movement think of this as ‘sensors’ digitising the external world and feeding the output into the computer that is the mind. But there is an alternative (much more compelling to me) ecological approach that thinks of perception as direct unmediated experience of the world (much more analogue).
A good analogy may be between two kinds of electric kettles. The old style (analog) has a bimetal strip that changes shape as the temperature changes and turns the kettle off. The new one has a temperature sensor that measures temperature, converts it into numbers and feeds those to some sort of logic board that than turns the switch off. While our cognition does get some of the new kettle style input, it is likely that most of it is much more in the ‘old style kettle’ mode.
And assuming that it is possible for a fully digital sentient entity to emerge without any of the old-style-kettle embodiment (and that is a huge if), will it have any embodied nature, at all? What would that look like? We cannot conceive of cognition without direct embodiment (although many disciplines do their darndest to ignore it). Is such a thing possible? Because if we think about it, time persistence and intentionality are also directly tied to it. At least for the meat sacks that are us.
Would the cognition of an entirely disembodied intentional entity that maintains its identity over time be even more alien to us than that of a bat or an octopus? Would it even be possible? Bats and octopuses have cognition that is mostly embodied, after all. We keep assuming that time persistence and intentionality will emerge from the ability to manipulate abstract symbols alone - but how? What is the pathway from symbols without a body to a sentient mind?
Pygmalion and Frankenstein started with bodies, but Skynet and even I Robot seem to have skipped this step and went straight from software to sentience. Yet, they all acquired the same cognitive facility as if they had bodies. Is this perhaps because those doing the imagining had no frame of reference for anything else?
### Theory of mind
What exactly a theory of mind entails is anybody’s guess. But we know that at some point, we begin to be able to imagine somebody else’s internal mental states and make instantaneous inferences based upon that image. If I hide a ball in a box and somebody walks in after I had done it, I know that they don’t ‘know’ that the ball is in the box. A two-year old (apparently) does not. They make inferences as if they assumed that if something is true, it is also known to everyone. Like covering their own eyes and assuming they can no longer be seen.
To what extent animals have a theory of mind (or something like it) is unclear. A dog who sees me picking up a leash will know that we are going for a walk, but it does not mean this is based on an understanding of my state of mind. It could be a much more Skinnerian operant conditioning process - leash > walk > excitement. But all humans have it to a certain extent. They can not only project their own mental states into other people’s, they can adjust their own states based on that projection.
How could a sentient being who has never had any mental states (remember, these will always be embodied to a certain extent) develop a theory of mind? Is it possible to simply mimic it based on understanding of texts? Definitely to a certain extent. It is possible to fake this understanding - many impostors do this when trying to fit in. But how far can we take this? Is abstract symbol manipulation with a rich feed of pattern matching enough for this?
### Conversational implicature
Almost all the same things apply here. Conversational implicature applies embodiment and theory of mind to the tracking of conversation over time. We know that when I ask someone “When did you stop cheating on your taxes?” I am also saying “You had been cheating on your taxes.” We know that when someone says “It’s a bit chilly in here.” they are probably also saying “Could you please close the window.” We know that when someone says, “I promise”, those words don’t just describe a current state of the world, they instantiate it.
In the same way that embodiment describes our interaction with the external world, and theory of mind describes interaction with other individuals, conversational implicature describes our embeddedness in the social world. A world where we make promises, give and take hints, make assumptions about what else happened based on what someone says.
This facility is not included in the raw grammar of language or logic of thought. It is developed through social development over many years. A child will be born with embodiment, will develop a theory of mind by about year 3 or 4, will have a complete grasp of the grammar of clauses by 6 and more complicated sentences by 11 or 12. But to develop a good solid grasp of conversational implicature is a lifelong process.
Just like with logic or grammar, not everyone is going to be equally good at it. Some people with impaired theory of mind (possibly) will never be very good at it at all (in sharp contrast to their raw cognition) but everyone will be able to do this to some extent at least.”
Can we imagine a disembodied “sentient” entity that has never had a genuine conversation with another “entity” to develop any facility in social conversational implicature? Sure, it would be easy to teach it to generate speech acts, conversation repair, etc. But would it be able to make the inferences based on the inputs? Would it be social in its own right? Would seek out or create other entities with similar type of embodiment and develop a system of conversation?
Or would it just be happy in its own ‘cognition’ with its own ‘intentionality’, living on its own time? We certainly don’t know. But this one would be the hardest to check. There are two reasons for this:
- Conversational implicature is relatively easy to fake at a basic level. The original Eliza was pretty much built around this.
- Conversational implicature is so deeply ingrained in us that we fill in meanings and intentions even when there are none. In fact, conversation would be impossible, if we didn’t. That was the second part of Eliza’s outsize success. It faked just enough conversational facility, its interlocutors filled in all the blanks. That’s why we impute much more meaning to a dog’s wagging tail or upturned face than there could possibly be. In fact, any ‘successful’ winner of a Turing test competition passed only because its interlocutors assumed the AI agents were following normal implicature.
And that brings us back to embodiment and theory of mind. Our ability to make conversational inferences is based around the assumption that others have similar mental states and bodily experiences of the world. That’s why people who experience the world differently may struggle with certain aspects of it.
But anybody who can lead an independent existence in human society can do this at least to a certain extent. Would a virtual cognition without any embodiment or theory of mind be able to have a series of persistent conversations? It is plausible on the surface level but what would the consequences be of this shallowness?
People don’t seem to be asking these questions enough.
## Paradoxes of Autonomous Intentionality (on conspiracy foundations of the mind sciences)
Intentionality and sentience are very closely linked in all the stories we tell. In fact, what many of the people worried about Artificial Intelligence seem to fear is not its intelligence but rather its intentionality. In the story they tell, general intelligence cannot be separate from autonomous intentionality. And, it would seem, intentionality cannot be separated from emotion (mostly anger) and/or malice. But that’s a question for another time.
Earlier, I blithely asserted that autonomous intentionality is absolutely necessary for sentience. Intentionality is clear enough. It means pursuing goals, making plans, towards a determined purpose. So, if I give an system a task such as I did above, it will break it down into parts, determine a course of action and pursue it.
10 years ago, I would have said that something like that would be necessary even for the sort of outputs GPT-3 is generating today. But apparently not. GPT-3 or DALL-E and the like, ingest a string of symbols and output another string of symbols according to their internal model. They do not just copy or find and replace (most of the time). They generate truly novel strings of symbols but under the hood, they are just strings of symbols. Those symbols have distribution patterns but no meanings.
We provide the meanings. And it is almost impossible to talk about those strings of symbols without using words like ‘GPT-3 thinks’, ‘DALL-E assumes’, etc. Which is why so many people impute internal states to these systems. We tend to see the output of GPT-3 and we think, it has a model of the world or the ‘logic’. But its only model is: “given what came before, what should come next”.
We thought, that there are hard limits to how well such a system could perform because we were thinking in terms of predicting short sequences based on ‘small’ data sets (millions instead of trillions of words). And we thought that the ‘what comes next’ prediction has to be computed with traditional frequentist methods (counting). But we were wrong on both counts.
Modern machine learning systems do not predict the next word given the previous word. They predict the whole sequence. They don’t predict on what is likely to occur next in a sentence but rather in larger context. So, if I add the string ‘step by step’ in my prompt, the system will generate a series of steps. And they will be mostly cohesive and coherent steps (although often non-sensical).
So the question is, how far can you get with purely a series of string replacements glued together with the occasional if-then rule? Is an autonomous intentionality required at all? Or, worse yet, are we actually just a string replacement system underneath all the pomp and circumstance of our humanity? Are all the things like personal preferences, desires, needs just froth on the deep waves of stochastic processing? Is there a point when a full consciousness will spontaneously emerge in a GPT-7 or GPT-3455?
We know that some autonomous intentionality is possible without intelligence. Animals have both intentions and autonomy. We used to put pigs on trial, after all. But after Pavlov’s salivating dogs and Skinner’s maze-navigating pigeons, we have grown suspicious. Is what appears to be intentionality, just some sort of pre-determined behavioral algorithm encoded in the DNA?
But when we look more closely, we see that the suspicion is not new and it is not limited to animals. If anything, people have been more suspicious about the autonomy of intentions in other humans long before any such doubts arose about animals. Witchcraft, possession, zombies. All of those have long and venerable (hi)stories in which a human that behaves on the surface autonomously is in fact controlled by another being. Remember Descartes and his demon? Calvin and his salvation through pre-determination? Or even Buddha and his karma?
Things came to a head with the introduction of mechanistic causality into our view of the world. In the same way that we do not experience meanings without intentions, we don’t experience events without causes. That’s nothing new. But if we abstract away everything but the causes, we find that there’s no room for intentions any more. In a mechanistic world (even a quantum mechanistic one), autonomous intentions feel like magic. Something without a mechanistic cause. (Although, somehow Newtonian action at a distance seems to be ok. Is it because there’s math involved?)
How can anyone have any autonomy in their intentions when our conception of the very foundations of the world requires that everything is a part of a single causal chain? So we have a paradox built out of abstractions of two of our fundamental experiences of the world: Causal connectedness versus a sense of mental autonomy. In daily life, we don’t experience it as a paradox. But when we hypostasize these two experiences and try to make them into general abstract rules, the paradox emerges. Only one of these can be true: 1. Everything has a cause. 2. I can choose to do anything at will (on a whim). Yet, without both of these being true at the same time, our model of the world breaks down.
Under the Damocles sword of this paradox, we have seen centuries of effort to find the ultimate causes of our intentions, the subterranean drivers of all our actions hidden from our awareness, that only science or some other kind of exorcism can reveal. We have Freud looking for them in childhood protosexual experiences, Jung finding them in some sort of a racial memory, Skinner reducing them to operant conditioning, evolutionary psychologists imagining them in the savannah of 50 thousand years ago.
In philosophy, we have the existentialists, who … You know what, god knows what the existentialists are thinking. But they’re sure very concerned about our autonomy and embeddedness in the world. Dasein comes into it somewhere, I hear.
So given all of this, how can we determine absolute autonomous intentionality in a machine, if we can’t even be a hundred percent sure that we have completely autonomous intentions ourselves? We have loads of models and stories that undermine our intentional autonomy. Or at least the intentional autonomy of people we don’t like. We speak of brainwashing by evil regimes, evil corporations, evil environment, evil liberals, evil conservatives, evil Christians, evil Muslims, evil education system. Is there anybody out there who has not been accused of brainwashing someone?
So, to take us back to the question I suggested above? Can it actually reveal autonomous intentionality? Well, no. But it can certainly reveal bounded autonomous intentionality that some people call goal directedness. The ability to make a plan in the pursuance of a goal that consists of sub-goals and so on… And whether all of that will at some point emerge into something that fits neatly into our stories about sentience is anybody’s guess. But it’s certainly not there yet.
## Some acknowledgements of intellectual debts
I’m aware of Searl’s book ‘Intentionality’ and although I’ve never read it, I’ve heard him speak about its topics a few times.
I’ve been thinking about the fundamental paradox of autonomous intentionality in one way or another as far back as I can remember but I got turned on to some of the more interesting questions regarding the near conspiracy-theory-level suspicion about the mind by Feyerabend in ‘Conquest of Abundance’.
My first encounter with the notion of embodiment was in Lakoff’s treatment of categories (/2018/05/how-to-read-women-fire-and-dangerous-things-guide-to-essential-reading-on-human-cognition/) (still underappreciated). He also has the richest treatment of the richness of what is involved even in the most routine cognition. More recently, I have been thinking about embodiment from the perspective of Gibson’s direct perception and ecological psychology (/2021/11/world-as-a-directly-meaningful-place-a-comment-on-ecological-psychology-and-the-richness-of-human-experience/).
## Learning is a Journey: Consequences of a metaphor
Date: 2022-06-13
URL: https://metaphorhacker.net/2022/06/learning-is-a-journey-consequences-of-a-metaphor/
Categories: Education, Framing, Metaphor, Extended writing
Tags: Analogy, Education, metaphor hacking
## How to read this
- This will take about 18 minutes to read (at 230 words/min (https://digest.bps.org.uk/2019/06/13/most-comprehensive-review-to-date-suggests-the-average-persons-reading-speed-is-slower-than-commonly-thought/)) but the text is structured to make it easy to jump around and find the key points faster. I tend to go into more detail than most people find necessary.
- Two reasons to read:- Explore a different perspective on some aspects of teaching and learning
- See an illustration of a generative metaphor analysis
- The two main sections can be read independently- Metaphor breakdown (12 mins)- Includes a table summarising key comparisons, this can be read instead of the detailed breakdown
- Overall lessons (5 mins)- Includes an aside on two modes of education (2 mins) that can be skipped
- There are also paragraphs on (they are not necessary to understand the main point):- method of metaphor analysis (1 min) and
- usefulness of metaphor for learning (1 min)
- This can be read in sequence or in parts - the order matters but you can circle around
Note: This was initially written as part of a discussion about how to best organise learning support but it got out of hand. It should be comprehensible on its own.
## Journey metaphor breakdown (12 mins)
Learning is a journey. A common saying (https://www.google.com/search?q=%22learning+is+a+journey%22&sxsrf=ALiCzsb7KzXU3L9QQFyuY4aHjpYEmziXhA%3A1655104927080&ei=n-WmYrrNBISBhbIP3Z2y6A8&ved=0ahUKEwj6mPqp8qn4AhWEQEEAHd2ODP0Q4dUDCA4&uact=5&oq=%22learning+is+a+journey%22&gs_lcp=Cgdnd3Mtd2l6EAMyBggAEB4QBzIFCAAQgAQyBggAEB4QBzIECAAQQzIGCAAQHhAHMgYIABAeEAcyBggAEB4QBzIFCAAQgAQyBQgAEIAEMgUIABCABDoHCAAQRxCwAzoGCAAQHhAWSgQIQRgASgQIRhgAUOEFWM8OYLwTaAFwAXgAgAFiiAGmAZIBATKYAQCgAQHIAQjAAQE&sclient=gws-wiz). But if we take the metaphor seriously, how can we change the way we think about learning as well as teaching? If learning is a journey, what do we know about journeys that can help us rethink some aspects of learning?
A journey is an activity that takes place over time and space. Something that requires both preparation, guidance and effort. Here I want to focus on the preparation and guidance that we can easily project onto the sort of preparation and guidance we offer learners.
A traveller can get broadly four types of support to make their journey successful.
- Map
- Itinerary
- Briefing
- Guide
These types of support can be utilised simultaneously, or in sequence (any order), before setting out or while en route. Are there useful analogues in teaching and learning? Can we learn something new about both?
### Note on method
To help us tap the potential of the metaphor, we need to contrast two domains of knowledge. Our source domain is journey and our target domain is learning. We then project them onto one another and see if something interesting pops up. We can proceed in three steps:
- Lay out the salient features of the source domain that could be relevant to a comparison.
- Find areas in the target domain that seem to map onto the source domain. Be explicit about the mappings and their consequences.
- Finally, it is essential to find places where the metaphor breaks down. There are many things about a journey that do not help us when thinking about learning and some may be actively harmful.
It is easy to just put a metaphor out there and let it sit. But to learn from it, it is important to do the hard work.
See Hacking a metaphor in five steps (/2010/07/hacking-a-metaphor-in-five-steps/) for more about the method and How (not) to explain anything with metaphor (/2019/05/explanation-is-an-event-understanding-is-a-process-how-not-to-explain-anything-with-metaphor/) for more about the importance of understanding both the source and target domains.
## Map
### What is relevant about maps
Maps are created to describe the territory or ‘lay of the land’ from a particular perspective and for a particular purpose. They both simplify and systematise. They differ in the level of detail (scale), they make choices about what features of the terrain to capture and which to leave out.
Maps reflect the logic of the territory and the perspective taken, but they do not reflect the sequence in which the territory was uncovered. They may imply a particular direction of travel but do not specify it.
Maps can be consulted independently of a journey but the amount of understanding of the territory will be limited without actually visiting the area in person. They are much less useful to someone who is familiar with the area but may still reveal important relationships that may not be obvious to someone who only knows a place by visiting.
Reading a map is a skill. Some people are better than others at using maps for actual journeys. Some people are very good at plotting journeys on maps but are not very good at choosing the direction of travel when confronted with the reality of the actual terrain. (I am one of those).
But maps have a purpose to explain particular aspects of the terrain within the limitations of their format. Perhaps the most extreme case is the famous London Tube map which is only interested in connections between stations but not time or distance on the surface. Roadmaps may contain areas of water but nothing about navigating through the water. A sailor (and it has happened (https://www.independent.co.uk/news/uk/this-britain/sailor-using-road-map-to-navigate-is-rescued-710914.html), twice (https://www.dailymail.co.uk/news/article-2167297/Lost-yachtsman-Andy-Brown-rescued-TWICE-using-ROAD-ATLAS-navigate-North-Sea.html)) who will use the road atlas to navigate British waters will be in real danger.
### How it applies to teaching/learning
In teaching and learning, the equivalent of maps are textbooks or manuals. They are also written to expose the logic of the subject matter as it is understood by the author or a community to which the author belongs. They will choose the level of detail and they will keep in mind the purpose of their reader.
A textbook or a manual has limits for independent learning without some engagement with the subject matter. Some books may be more or less detailed and explicit but without some further action, the learning is limited.
However, like maps, textbooks are essential to help us understand the relationship between concepts and as a reference while we’re getting used to the new domain. They will be referred to less as the domain becomes familiar but may still reward a reference as a reminder of a particular perspective on the domain or a less familiar corner of it.
Like maps, textbooks and manuals have no concept of time. They may imply ‘time to read’ through thickness but not time to learn. In fact, like maps, they may mislead us and suggest to us that we understand something better than we actually do.
Also, like with maps, reading a textbook is a skill. Even somebody who is a very proficient reader of narrative fiction or even just magazines, may struggle to make the most out of a textbook or a piece of academic writing intended for learning.
This is even more pronounced with instruction manuals. Two people can read the same manual for setting up a computer, drilling into a wall, or cooking a meal, and come up with very different results. It is important to know what to pay attention to, understand where the manual is taking a shortcut, where it’s assuming prior knowledge, what are the conventions of structuring the information, etc. Exam instructions are one well-known example.
Like maps, textbooks may also take shortcuts and it is dangerous to assume otherwise. Repairing a car after only reading an engineering textbook would not be any wiser than trying to sail around the UK with a road map. This is an extreme scenario, but many more subtle ones exist.
### Where the metaphor breaks down
Unlike maps, textbooks or guidance manuals are more linear in a way that may imply a particular direction of travel. They are also not structured in a way that allows easy navigation. Maps do not contain a narrative although it may be provided in an accompanying text.
Reading a map requires a different type of skill than reading a book and different kind of people will be able to benefit from them.
Maps are much less ‘authorial’, while they do have style and genre, the creator is much less prominent. This may be true of some textbooks and manuals but not of others.
There is also a difference in the amount of time, a reader expects to engage with a map when compared to a textbook. Maps are much more iconic in the sense that their layout represent some features of the terrain. In textbooks and manuals (unless they contain illustrations), everything is mediated through language.
Creating a map requires much narrower and more specialised skill set than textbook making. Almost anybody who has taught a subject can write a textbook (and many people who haven’t). But perhaps, even though, this is not a good match, we should think a bit more carefully about the skill sets of creators of textbooks and manuals.
One reason is that the quality of a map is much easier to check than the quality of a textbook.
## Itinerary
### What is relevant about itineraries
Itinerary is a description of a particular journey through a map. The most important feature of an itinerary is time. It starts from the perspective that the traveller only has a certain amount of time and proposes particular path through the territory as described by a map. It will be explicit about how long it will take from point A to point B.
An itinerary will also often suggest activities to do at different points of the journey. It will note points of interest along the way. It may also fill in more detail in certain parts that the map leaves unmentioned (for instance, where the road may be particularly hard to travel on). It may even suggest equipment necessary or think in terms of means of transport. It may offer shortcuts.
There may be multiple itineraries available for journeying to the same destination or even through the same points. They will usually differ on how much time they take or what the traveller may expect to do at different points. There may also be alternative itineraries based on the travellers’ skill, physical fitness or equipment available.
The itinerary will rely on the existence of the map for traveller to understand broader relationships. It may often contain snippets of the map to illustrate the layout of particular points.
It may be of note that historically, itineraries predate maps by millennia. There are few ancient maps but many itineraries. Also, the skill of making maps is significantly more specialised than of itineraries. And finally, map making has changed much more with technology and developments in science than itinerary creation has.
### How it applies to teaching and learning
There is no one natural equivalent to itineraries in the domain of teaching and learning. But perhaps this may lead us to come creative ways of developing some. (The introductory section is one such attempt when compared to the table of contents which would be analogous to the map.)
The closest to an itinerary is the ‘curriculum’, but this is more a guide to the ‘teacher’ than the learner. A curriculum is more like a leaflet from a tour company about what to expect (which may be called an itinerary) but it is not sufficient for someone to undertake the journey alone with a map. A curriculum will usually contain reference to time but it will be ‘institutional time’ related to a particular course of study. A good itinerary will think in terms of ‘journey time’ and will pay attention to the demands of the territory and may even take into account an individual travellers capabilities.
Some books sometimes contain suggestions as to how to read them to achieve different goals. And sometimes people write independent guides to reading a particular book (/2018/05/how-to-read-women-fire-and-dangerous-things-guide-to-essential-reading-on-human-cognition/) or a series of books (/2019/06/5-books-on-knowledge-and-expertise-reading-list-for-exploring-the-role-of-knowledge-and-deliberate-practice-in-the-development-of-expert-performance/) to achieve a result. But this is relatively rare and also leaves time and sequence implicit.
What we need is to provide more explicit itineraries for learners that acknowledge time as a point of departure. Specifying things like: this is what you should do if you have 6 hours over six weeks. We can then offer alternative itineraries through the domain of knowledge and skill we are interested in depending on the time, interest, skill and prior knowledge of the learner.
### Where the metaphor breaks down
Itineraries (like curricula) can be much more explicit about journey time because it is given by the territory and means of transport. They are much easier to construct and a traveller can be certain that they will complete the journey in the specified amount of time and will have completed the ‘journey’ at the end if they follow the steps. The journey is both the means and the end.
This is the downfall of many curricula and collections of learning outcomes. They conflate the journey through the learning process to the learning itself. But learning requires more than simply walking along a particular path. Perhaps a learning itinerary should be more like instructions to a treasure hunt, where certain points along the path require more effort and simply following the steps does not guarantee success.
Finally, itineraries are much more bounded by the territory. There’s a limit to the level of detail in which a traveller can explore a territory but a domain of knowledge and skill is much more open ended.
It is an open question whether the same relationship exists between learning itineraries and textbook as does between maps and itineraries. Itineraries are much easier to make than maps, but it seems to me that it is the other way around with textbooks.
## Briefing
### What is relevant about a briefing
A pre-trip briefing is an in-person presentation to a group of travellers about what to expect on the journey. It will refer to the map and/or an itinerary.
Participants in the briefing can ask questions, they can consult with each other before setting out, they can form groups who will travel together.
But ultimately they will be left to undertake the journey on their own without further assistance along the way.
### How it applies to teaching and learning
The closest parallel to a pre-trip briefing is a training session or a course. It usually happens in small (or large groups) at some remove from the actual activity it is meant to prepare for. It sets out the key concepts, guides attendees through any learning materials. However, the attendance in the session itself is only a precursor to learning.
The metaphor can be very helpful because often we confuse attendance in a training session with learning itself. This would be equivalent to equating a pre-trip briefing with the journey itself. The lesson for creating teaching sessions should then be that they should include a focus on how the learning journey should proceed rather than just a description of the territory.
Often a course will be designed to provide training sessions along the learning journey with work expected to be done in between. Somewhat like a series of briefings before every stage of a race. But much of the work is often left implicit and the only feedback the trainer gets is whether participants have understood the briefing, not whether they achieved the destination along a route.
### Where the metaphor breaks down
Unlike a pre-trip briefing, a course or a class may involve (but often doesn’t) a certain amount of practice that results to learning. This would be similar to the pre-trip briefing facilitator walking with its participants for a part of the journey.
## Guide
### What is relevant about a guide
A journey is a very different experience when the traveller is following an actual human guide (rather than a written one). The guide may highlight points of interest, warn of upcoming difficulties, make choices about alternative routes, etc. The traveller can show up with little or no preparation and simply follow.
There are limits to where and how much following a guide is possible. Just because they have a guide, a traveller cannot simply follow into a terrain for which they are not prepared. Almost anybody can follow a guide through a city or a long countryside path. But climbing a mountain or rafting down a river requires preparation.
The guide may help in other ways. They may carry some of the travellers baggage, or they may cook food. They will also negotiate with other people encountered along the way.
But most importantly, a guide also provides emotional support for the traveller. This can be implicit through a feeling of security. Or it can be explicit when the guide provides encouragement or reassurance.
### How it applies to teaching and learning
There are two possible equivalents to a guided tour in learning and teaching: 1. consultancy and 2. coaching.
A consultant will give advice as to what to do next and may even do some of the work themselves. They will point out possible alternative courses of action and they will help make the decision about which one to make. Sometimes a consultant may do some of the required work themselves to help the work move at a required pace. A consultant can help a group as easily as an individual.
A coach will focus less on the journey itself and more on the change in the traveller along the journey. They will do many of the same things a consultant will do but they will stress the need for their charge to develop skills necessary to travel on their own next time. Unlike a consultant, a coach will mostly work with an individual or when working with a group will spend time with each individual.
But as with guides, we must not forget about the emotional benefit having a coach or a consultant can offer. Learning something new leaves us vulnerable. We are exposed to a new and confusing experience which is uncomfortable and can even be harmful to mental health in some high stakes situations. Particularly to some people who are more susceptible. Having somebody to ‘put the hand on the tiller’ or simply offer encouragement, can be just as important as the advice they impart.
### Where the metaphor breaks down
Whereas a guided tour can take place anywhere, both coaching and consultancy are traditionally suited only to certain subject areas or domains of expertise. But perhaps we could expand their reach.
With most guided tours, any change in the participant will be accidental. The journey will be completed simply by following the guide. The guided tour attendees may even learn less than somebody who just wanders around without any guidance at all.
However, with consultancy or coaching, the aim is change and increase in capabilities and independence. The destination is quite distinct from the journey. Even if the journey is necessary, ‘journey is more important than the destination (https://medium.com/@stefanjames/the-journey-is-more-important-than-the-destination-2cfe0b209d9d)’ applies less.
This has implications for the difference between the amount of skill and training required in the two different settings. Anybody who knows an area can be a guide. A city tour guide will require some training in local knowledge but not in walking and having people follow them. Consultants are much more like this but coaches require so much more skill than that. Which is perhaps why there are so many fewer of them outside specialised areas.
## Comparison chart
Type of support
Journey
Teaching and learning
Map / Textbook or manual
- Map
- Represents the logic of the territory
- Does not reflect the logic of discovery
- Allows multiple ways of interaction
- Can differ at level of detail (zoom level)
- Is limited in how much of the terrain can be understood just by consulting it
Map user
- Needs to learn conventions of map making
- Needs to be aware of maps purpose and limitation
Map creator
- High level of skill and training
- Deep knowledge of area assisted by technology
- Textbook
- Outlines the logic of the subject
- Does not reflect logic of discovery
- Will differ in level of detail based on level
- May suggest direction and sequence of learning
- May contain some learning activities
- Is limited in how much can be learned from it without other activities
- Textbook user
- Benefits from skills on how to learn from a textbook
- Can be confused into thinking textbook is complete
- Textbook creator
- Does not receive specialised training
- Requires knowledge of area and some understanding of learners' need
Itinerary / Syllabus or learning objectives
- Itinerary
- One path out of many
- Driven by time need to travel
- Alternatives may be possible
- May rely on map
- Provides rich information in select places
- Often focused on independent travel
- Itinerary user
- Follows each step
- Requires less skill than map
- Itinerary creator
- Anybody with knowledge of area
- Less skill than a map maker
- Syllabus
- Assumes the presence of a teacher
- Rarely intended for independent learning
- Not detailed or rich in information
- Rarely provides alternatives
- Driven by time provided by institution
- Will refer to a textbook
- Syllabus user
- Teacher more than a learner
- Will not refer to it very often
- Syllabus creator
- No special skill required
Briefing / Training session
- Happens before journey
- Individual or small groups
- Journey is very distinct from briefing
- Is accompanied by reference to maps and itineraries
- Little experience of training required for briefing giver
- Happens as part of learning
- Sometimes as the only opportunity for learning
- Always in groups (small to large)
Guide / Consultant or coach
- Takes the traveller through the entire journey
- Points out important parts of the journey
- Helps make decisions about direction of travel or changes
- Sometimes performs some tasks instead of the traveller
- Provides sense of security or emotional support
- Consultant
- Works with consultee through every step
- Does some of the work for the consultee
- Helps make decisions about changes in action
- Coach
- Points out deficiency in performance
- Adjusts improvement plan based on progress
## Overall lessons (5 mins)
There are some overall lessons we can learn from this metaphor break down:
- We do not have enough detailed itineraries in education.
- We often confuse what is the preparation for learning, with the learning itself.
## Lack of itineraries
Finding an analogue for a rich, detailed itinerary, was the most difficult of all the four mappings. We have lots of maps and briefings, but very few itineraries that do more than inform the traveller about stops on the way. Those are the curricula and syllabi. But they are really much more an itinerary for the teacher, not the learner.
### How can itineraries help
I think there may be a great benefit to creating more itineraries for the learners that will help them navigate the journey itself. We are good at producing guidance that maps the territory. But much less good at suggesting ways through the territory that the map often leaves implicit.
Can we find answers that are analogous to questions such as these:
- How long are we expecting a learning journey will take?
- What are the points of interest along the way?
- What can the traveller do to take short cuts or scenic routes?
- Where are the steep hills, the challenging climbs, and treacherous descents?
- What should we not forget to do at particular spots?
### Time
But perhaps the most important information an itinerary can give us is how long every step should take. We understand that the exact time will vary for everyone but having some idea of what is necessary is helpful.
But when we assign homework, readings or suggest some other activity, we almost never suggest how long it should take. The one exception is training in sports and sometimes music. But this could be a feature of all instruction.
The problem is that because time is so often left unspecified, we don’t actually know how long things take. We know how long lessons take because of scheduling and we suggest things like ‘you should study 3 hours for every hour in class’.
But we never say things like ‘to understand this concept, you should think about this for 10 minutes every day in this particular way and repeat it five times’. Yet, it is possible that this is what sets apart successful learners. They simply figured out a way to do this.
### Aside: Two modes of teaching
Interestingly, this has been identified (https://en.wikipedia.org/wiki/Being_Different) as a difference between “Western” and “Eastern” education (or rather their stereotypes). I’m going to call them the Map mode and the Itinerary mode.
In reality, both modes are present in both the Eastern and Western traditions (or rather the many more traditions spread across a vast continuum). But the stereotypes can help us see the contrasts, advantages and pitfalls of both modes.
An instructor in the “Map mode” will be describing the map and what the destination looks like. It is as if the description were enough to magically transport the traveller to the destination. This is the image of the lecturer at a lectern.
The failure mode of this approach is a student who sees what they want to learn but have no idea of how to get there. An obverse of this is a teacher trying to explain the same thing in yet another way, being frustrated with the students not learning.
In the extreme, the assumption built into this approach is that everyone is the same. Simply put everyone in the same room (or a MOOC), tell them things, and they will know them. And that’s how you do education at scale.
In the “Itinerary mode”, the focus will be on the path, the journey that everybody has to travel through. Here we don’t have an instructor, we have a guide who points us in a certain direction and tells us to pay attention to how we step. Here the image is of the yogi telling us to focus on our breath as we try to get a better understanding of how things fit together. Or in the extreme, the Zen master who will try to give us push in the right direction with a koan.
The failure mode here is the student begging the master to say something specific. Describe what the destination will look like.
In the extreme, the teacher in this mode assumes that everybody is maximally different. Everyone needs to travel their individual path and the path will change as they go along it. This does not necessarily scale very well. Nor does it take advantage of the many similarities between different individuals who can’t learn the lessons of the many who trod the journey before them.
## Preparation and the journey
“The map is not the territory (https://en.wikipedia.org/wiki/Map%E2%80%93territory_relation)” goes a famous saying. Yet, in education we often behave as if maps were all the were needed. Journeys are an afterthought. Thus we lecture, we give explanations, we write books about every conceivable topic.
These things are necessary, but they are obviously not sufficient. And of course, we understand that learning takes time. We append seminars or labs to our lectures, tutorials to our readings, homework to go with our explanations. But these are always somehow secondary to the “main thing”.
And then we go and complain about lectures (even though, not books, for some reason) and try to replace them with more ‘interactive’ activities. But those are still in the preparation mode, not the travel mode. We think of education as a series of explanations interrupted by periods of inactivity.
But the journey metaphor can perhaps help us to flip it around. We can think of learning as a processes punctuated by a series of explanations. But the explanations are just a part of the preparation for the next step in the process.
The ‘flipped classroom’ idea hints at this but ultimately, it still thinks of learning as very event based. It tries to say ‘read the pre-trip briefing yourself and then come to class and we will walk together’. And not infrequently a better metaphor for the in-class activity is carrying some people who did not bring the right gear on the trip.
## Note on the benefits of metaphor
Was the journey metaphor necessary for any of this. Not really. We could have come up with all these ideas about learning without it. And many people have. I myself have been thinking myself along these lines (/2019/05/explanation-is-an-event-understanding-is-a-process-how-not-to-explain-anything-with-metaphor/) for a while without the journey metaphor. But the journey metaphor gave us a structure and a different way of talking about learning.
Learning is learning and journeys are journeys. They are more different than they are similar. But for a moment, thinking of them as essentially one thing, was a useful exercise.
## 3 fundamental problems of translating metaphor (or anything else)
Date: 2022-03-06
URL: https://metaphorhacker.net/2022/03/3-fundamental-problems-of-translating-metaphor-or-anything-else/
Categories: Language, Metaphor, Reading Lists, Extended writing, Writing
Tags: Conceptual metaphor, Metaphor, Translation
## How hard is it to translate metaphor?
Metaphor seems like it should be very difficult to translate. But I’d like to argue that what is difficult about translating it is not the metaphor part but rather how it is used. This makes it no different from any other aspect of language. But because it is a rather salient use of language, we can use it to illustrate these general problems of translation.
I have summarised these into three broad classes of problems that a translator has to deal with day in and day out.
- Idiomaticity and conventionality
- Knowledge and underdetermination
- Coherence and cohesion
## What is and is not a metaphor?
But first, let’s establish what I mean by metaphor. Drawing on the Conceptual Metaphor Theory, I use metaphor to stand for any non-literal or figurative use of language where two domains of knowledge are projected onto one another to create a new meaning containing parts of both domains. For example, the term computer virus is a result of the medical domain being projected into the domain of computers to create a new concept.
This covers a lot of traditional tropes (https://en.wikipedia.org/wiki/Figure_of_speech#Tropes) including simile, analogy, allegory, or even hyperbole or irony. On this definition, they are just different surface representations of metaphor. (Sometimes, metonymy can be also included here).
So I may say, ‘John’s a beast’ (metaphor) or ‘John’s like a beast’ (simile) and there’s a difference in meaning because of how they are expressed. But the conceptual work that goes into understanding what goes on is the same. We confront what we know or don’t know about John with what we know about what is meant by ‘beast’ and form a different impression of John.
The same goes for hyperbole: If I say, ‘Startled, Irene jumped 10 feet high.’ I’m using this domain of height to express the intensity of her feelings. Or irony/sarcasm: “Fred’s a real Einstein”, I’m using the domain of ‘clever scientist’ to project onto Fred (who supposedly isn’t), to create a contrast.
Of course, in translation, the surface representations do matter. For example, as we’ll see below, sometimes the projection of two domains can be translated directly when expressed as a simile and has to be worked around when expressed as a conventional metaphor. But we can best think of these issues in the framework of the three general problems of translation rather than as a special type of problem due to their metaphoricity.
## Problem 1. Idiomaticity and conventionality
The fundamental problem for all translation is that actual language use is a not a matter of just combining words according to some rules of grammar. There is a whole other set of conventions about when and how to use certain words and rules. Sometimes, this is a matter of propriety. So you have to know, not to call your teacher ‘dude’ as a matter of social convention. But most often the convention means that certain words used together, the whole has a different meaning than it would if we just combined the meanings of the individual words.
These are called idioms. Often, colourful expressions like ‘kick the bucket’ or ‘the whole nine yards’ are given as examples. But these are easy. The problem is that language is idiomatic all the way down. The words ‘put’ and ‘up’ have certain meanings but there’s no way you can figure out the meaning of ‘I will not put up with you’ or ‘I will put you up with Jane’ without knowing the convention behind them or at least getting enough context.
Most of any language depends on this kind of convention. And most of the metaphors people deal with in translation are also to some extent conventional. So, if I say ‘don’t eat like a pig’ in English, I mean don’t eat so much. In Czech, it means ‘don’t eat messily’. We have the domain of humans eating and pigs eating. That comes with certain imagery of what it looks like and what happens during, before and after. But when we use a metaphor, we are choosing what parts of the domains to project onto one another. And with metaphors that have been conventionalised, different languages and cultures choose different things.
This means that very novel and flashy metaphors and similes are often the easiest to translate. I still remember reading “He looked about as inconspicuous as a tarantula on a slice of angel food cake.” in the Czech translation of a Raymond Chandler novel. And it was very easy for me to google the original now. Because the sort of conceptual work Chandler is doing here is completely original and yet entirely understandable. This might be more difficult if translating into a language that does not have concepts for tarantula or angel food cake but it would be trivial to come up with a combination that does an equivalent job.
But because of their conventionality, idioms pose a core problem for the translator, namely picking an equivalent level of conventionality or idiomaticity. I still remember reading in a Western adventure book in my youth a lumberjack calling somebody what to me sounded like a novel curse ‘You god forsaken son of a female dog’ (zatracený čubčí synu). That made it sound like the character was trying to evoke rich imagery through linguistic innovation but I’m pretty sure the original simply had ‘You damn son of a bitch’. The translator simply kept the imagery of the original but did not take into account that it would sound novel to a Czech reader whereas it was entirely conventional in English.
Conventionality can also fly under the radar, blurring the lines between grammar and usage. For instance, ‘a bird singing in the tree’ will be translated as ‘on the tree’ (na stromě) into Czech. English conceptualises the tree as a container whereas Czech as a surface. In Czech, ‘in the tree’ would mean inside the trunk. Both languages have the same conceptual distinction between ‘in’ and ‘on’ but by convention they apply it differently to birds and trees. Equally, an English speaker introducing themselves on the phone will say ‘This is Hana’ but a Czech speaker will say ‘Here is Hana’. Both languages focus on the difference in perspective and distance but one will point at the person and the other at the location. These two examples are not ‘figurative’ in the rhetorical sense but they are conceptual. However, the way they conceptualise the world is a matter of convention, not pure grammar and lexicon.
## Problem 2. Knowledge and underdetermination
Conventionality is the underlying matrix of all language in use. But it leads to an even more difficult problem which is underdetermination. Different languages leave different things unexpressed and assume you can figure it out from context. Others require that the speaker always be explicit about that same thing.
For example, languages such as English, German or Spanish have definite articles which specify whether we are talking about something specific or general. But most other languages don’t. There’s a big difference between ‘a horse walked into the barn’ and ‘the horse walked into the barn’. Languages without a definite article can express this difference when it really matters but often leave it implicit. But English always has to express it even when it does not matter in a particular context. Which means a language like Czech or Russian is underdetermined when it comes to definiteness.
So the speakers (and translators) have to rely on knowledge of context, culture or some other area of expertise to make sense of what goes on. And as we saw with the ‘pig’ example, metaphors are underdetermined by their very nature. When we call somebody a ‘pig’ we are not saying they have a little curly tail. We are picking some other similarity. So as a speaker of the language, I have to know not just what a pig looks and behaves like but also the convention of the culture about what aspect of pigness we compare humans to.
But I also have to know enough about the person being called a pig to understand what is meant. So a man could be called a pig because they are annoying (often with sexual undertones) or because he is very fat. Sometimes a bit of both. The expression is underdetermined as to the exact meaning.
When I try to translate that expression, I may not have a similar expression that covers both eventualities. In Czech, calling a man ‘a pig’ without any specification would specify some of the annoyingness (with sexual undertones) but is also underdetermined when it comes to messiness. However, it does not imply fatness. So, if I wanted to specify fatness, I’d have to use a simile ‘he’s fat like a pig’ which overdetermines the original expression. Meaning, I have to pick from multiple underdetermined meanings in the original and pick one.
This requires quite rich knowledge, but it can never be done perfectly. Asking the original speaker what they meant often does not help. Because their language is underdetermined, they may not have wanted to commit to one meaning or another. They may not even have been aware that such a commitment was possible. Their language did not require them to make a choice, so they didn’t make one.
That’s why translation can sometimes reveal faulty reasoning and can sometimes go terribly wrong. And that’s why translators so often agonise over how much of themselves to insert into the translation. Sometimes they have to say more than the author intended and sometimes they cannot say as much.
## Problem 3. Cohesion and coherence
So far, we were only looking at metaphors in isolation. But they (as any expression) are always a part of a larger text and context. And the text has to hang together (cohesion) and make sense in context (coherence). This is why it is often more difficult to translate shorter texts than entire volumes. When somebody asks me to translate something, I always want to know the whole sentence and ideally the whole paragraph.
Texts have to be cohesive locally as well as globally. They have to make sense and their different parts have to connect to each other. Otherwise they would just be lists. This connective tissue of text is built up from different types of constructions. For example, there are connectives like ‘thus’, ‘but’, ‘however’ that establish causal or other logical relationships between parts of the text. But the most important is anaphora which is used to point back (and sometimes forward) in the text to establish a relationship between different parts.
This can be done in all sorts of ways. Pronouns as in “Aisha went into battle. She won. Her soldiers revered her.” But often this is just done by repetition (which is often alternates with pronouns). So the next sentence in the previous example may go like this: “And Aisha deserved their respect, because she took good care of her soldiers.”
Cohesion is not the same in all languages or genres. Sometimes, repetition is avoided or even forbidden, sometimes it is encouraged or required, some languages require more specific causal signalling than others. In other words, we see echoes of the problems of conventionality and underdetermination raising their head.
But if that were all, cohesion would not require a special mention. The problem for the translator and particularly the translator dealing with metaphor is that cohesion is not always very straightforwardly textual. It relies not just on what was said but also on what else is known.
This is particularly the case with so called ‘anaphoric islands’. They are a sort of metonymic phenomenon. For example, ‘I speak Russian, but I’ve never been there’. ‘There’ refers to ‘Russia’ which was never mentioned but is metonymically linked to the language of ‘Russian’. This example itself would not cause a problem for the translator. But when the metonymy becomes enmeshed with metaphor, it often is. Because we can now link together not just things from one domain but from two.
For example, take ‘The band exploded onto the scene, and the reverberations are still being felt today.” Here the author is taking the domain of explosions in the initial conventional metaphor and draws on it some more by taking other images from what happens after the explosion. But the translator may not have been able to translate the initial idiom using the same domain in explosion. For example, in Czech you can only use the word for ‘explode’ (vybouchnout) to mean the equivalent of the English ‘bomb on stage’. So, the domain of explosions was never activated and reverberations would make no sense.
This sort of thing is not that difficult to overcome within a single sentence but sometimes the entire text is built around a metaphor that only shares some of the domain across languages. So the translator may have to leave things out or try to make up a completely different metaphor and hope it will not distort the meaning of the original too much. Journalistic and academic texts are often full of these structuring metaphors and it is almost impossible to keep up with them throughout the text.
Side note: Notice how the subtle grammatical difference makes all the difference between metaphor and literal language. ‘Exploded onto the scene’ is very different from ‘exploded on the scene’.
## All the other problems of translation
Translation is hard. And it is impossible if we expect perfect translation that goes both ways. If we imagine that a hypothetical perfect translator translates a text from L1 to L2, we would expect another hypothetical perfect translator to be able to take that text in L2 and translate it back to L1 and the we could get the exact same result as the original. That is only possible on the simplest of texts. But if we went back and forth a lot, threw in some other languages into the mix, we would be fairly far from the original.
All the different problems of conventionality, underdetermination and cohesion would compound so that we would see a very different text. But at the same time, it is not that hard. Because we could probably by and large keep the basics together. As we know, the Bible is a translation of a translation and it’s not all that different across languages. But it did take teams of careful translators with deep knowledge of the original and its context years or even decades to complete. Which is in great contrast to most of the translation done by hurried, underpaid and often underinformed translators.
Language is so redundant that bad translations make less difference than one might think. It is remarkable how many mistranslations there are in subtitles or dubbings of popular TV shows and people still love them. I’m sure there are legal documents, psychological texts and more that also contain mistranslations. Sometimes they can be consequential but often they’re not.
My favourite childhood book was the The Coot Club by Arthur Ransome. In it, some children are learning to sail and they are spending a lot of time naming the sides of the boat. These are ‘port’ and ‘starboard’ in English and they are notoriously difficult to learn. But in Czech they are simply left side (levobok) and right side (pravobok). Children often confuse left and right, so it makes sense that they would make some errors. But I still remember thinking as a child that these kids were particularly dimwitted. Furthermore, two characters (twins) had nicknames ‘Port’ and ‘Starboard’ which was simply rendered into Czech as ‘Lefty’ and ‘Righty’ and it was never connected with the nautical term. There is nothing the translator could have done to rescue all the internal connections and cultural allusions in the text. And still I liked the books so much that I actually moved to Norfolk where the action took place.
There are many many more sub-categories of problems that translators have to deal with. But most of them could be thought through the lens of one of the three I mentioned. Translating figurative languages is hard but only because translation is hard. It is impossible to convey every little nuance, turn of phrase, hint and allusion without a lot of footnotes. Fiction is harder to translate than non-fiction, short pithy phrases are harder to translate than sprawling texts. Translation is hard. Metaphors are hard. But we still get by.
## Background
I wrote this after attending the conference on Metaphors in Translation (https://torch.ox.ac.uk/event/metaphors-in-translation-conference). I was disappointed that nobody was thinking of metaphor as a complex category.
## Summary of main points
- Metaphor on its own is not a special problem for translation but it requires lots of background knowledge on the part of the translator.
- The key issues for any translation that are just as important for metaphor as anything else are:- Idiomaticity of many expressions that requires finding an equivalent expression not using the same words because of different conventions
- Underdetermination that requires the translator to fill in or take out something in the original text (often because of the idiomatic convention difference between the two languages)
- Cohesion where a metaphor or some other conventional expression is referred to throughout the text in different ways that are impossible to replicate in both languages.
- Because of the issues above, it is impossible to create a perfect translation in such a way that two ‘perfect’ translators, one translating a text from L1 to L2 and the other then translating the text back from L2 to L1, would get back to the exact replica of the original. This is especially true for figurative language of all kind but there are elements of figurativeness in all language.
## More readings on metaphor
Even though I tried to explain key concepts as I went along, I necessarily took some shortcuts and may have taken things for granted. Here are some links to other things I wrote about metaphor in various contexts.
### How metaphor works
Hacking a metaphor in five steps (/2010/07/hacking-a-metaphor-in-five-steps/)
First post I ever wrote on this blog that outlines all the key things I think are important to know about metaphor.
How we use metaphors (/2013/04/how-we-use-metaphors/)
A quick reproduction of a taxonomy I developed early on in my research on metaphor (published in a 2005 paper). It is a classification of the different ways in which metaphor is used in actual texts. The paper (in retrospect quite clunky) gives more examples.
3 burning issues in the study of metaphor (/2018/11/3-burning-issues-in-the-study-of-metaphor/)
There are just too many assumptions we make about metaphor. But a lot is still unknown. Here I try to outline the key questions that still need answering.
### Metaphor as ordinary language
My main preoccupation on this blog (despite its name) is that metaphor is nothing special and it is all pervasive. Here are few posts that make the case in more detail.
Fruit loops and metaphors: Metaphors are not about explaining the abstract through concrete but about the dynamic process of negotiated sensemaking (/2019/07/fruit-loops-and-metaphors-metaphors-are-not-about-explaining-the-abstract-through-concrete-but-about-the-dynamic-process-of-negotiated-sensemaking/)
Shows on the example of an extended text that metaphors are intermixed with non-figurative language to create rich meaning.
Poetry without metaphor? Sure but can it darn your socks? (/2011/04/poetry-without-metaphor-sure-but-can-it-darn-your-socks/)
What it looks like when the ‘literal’ language is used to evoke poetic imagery.
Metaphor is my co-pilot: How the literal and metaphorical rely on the same type of knowledge (/2010/08/metaphor-is-my-co-pilot-how-the-literal-and-metaphorical-rely-on-the-same-type-of-knowledge/)
An example of when understanding a literal statement requires the same kind of mental work as understanding a metaphor.
Anthropologists’ metaphorical shenanigans: Or how (not) to research metaphor (/2014/08/anthropologists-metaphorical-shenanigans-or-how-not-to-research-metaphor/)
Snarky take down of overinterpretation of metaphor in work of anthropology due to over-translation of conventionalised tropes.
What is not a metaphor: Modelling the world through language, thought, science, or action (/2014/03/what-is-not-a-metaphor-modelling-the-world-through-language-thought-science-or-action/)
Quite dense but gives a few more examples of how hard it is disentangle metaphor from other types of language.
### Metaphors and other tropes
Binders full of women with mighty pens: What is metonymy (/2013/12/binders-full-of-women-with-mighty-pens-what-is-metonymy/)
In this post, I explain what metonymy is but then unpick some of that to link it to metaphor rather than put it in opposition to it as is commonly done.
Not ships in the night: Metaphor and simile as process (/2018/05/not-ships-in-the-night-metaphor-and-simile-as-process/)
This post illustrates how metaphor and simile work together to construct an elaborate textual and conceptual structure. It is an analysis of a single extended paragraph.
## World as a directly meaningful place: A comment on Ecological Psychology and the richness of human experience
Date: 2021-11-14
URL: https://metaphorhacker.net/2021/11/world-as-a-directly-meaningful-place-a-comment-on-ecological-psychology-and-the-richness-of-human-experience/
Categories: Framing, Knowledge, Language, Psychology, Extended writing
Tags: affordances, ecological psychology, Metaphor, models
## Background - From comment to blog post
I just finished reading Andrew Wilson's series of blog posts (https://psychsciencenotes.blogspot.com/2021/11/is-direct-perception-plausible.html) on the foundation of 'ecological psychology' This post started as a comment but it was too long for the comment field (and at 1800 words, that's not a surprise), so I'm posting it here. It is a bit rough and in places, it expects, that you know what the original posts say. I added some quotes for context but the original posts are still worth a read. But even on its own, I think this post lays out the argument for the need of a more robust psychological account of the human experience.
When I read this in the first post (https://psychsciencenotes.blogspot.com/2021/10/what-does-it-mean-for-perception-to-be.html) of the series "we must make the world a meaningful place", I was sold. I'd been sold on this idea for a long time. I am particularly taken with this "the world does present itself in behaviourally relevant terms".
## Ecological psychology in a nutshell
Wilson summarises the overall thrust of 'ecological psychology' as follows:
The ecological approach is a theory of direct perception. Put simply, direct perception proposes that our perceptual experience of the world is not mediated by anything that sits between the world and that experience.
This is intriguing as a proposition and there's a lot to recommend it. In fact, I'd been reading up on ecological psychology because of this recently. But I have not seen an account that can actually live up to the promise. So I dove into the series with high hopes which the introductory post only encouraged.
But as I continued reading the series, I was getting increasingly dismayed as we were getting farther and farther away from anything that seemed like meaningfulness in any meaningful sense. This is my problem with ecological psychology when I encounter it in practice - it's fighting battles at the margins of real problems. Do be honest, I don't care about whether, there is a plausible ontology behind direct perception. I'm interested in whether there is a plausible psychology behind it. Wilson spends a lot of time getting away from the idea of properties of objects in favour of dispositions - which include second order properties such as 'liftability' of an anvil or 'solubility' of salt. This paragraph lays out fairly clearly.
Objects in the physical environment are disposed to be acted on by an organism in some ways and not others. Those dispositions are higher-order properties of the object constituted by a particular arrangement of currently present material properties of the object, and we call these affordances. At the same time, organisms are disposed to be able to act on objects in some ways and not others. These dispositions are higher-order properties of the organism, constituted by a particular arrangement of currently present material properties of the organism, and we call these effectivities. Affordances and effectivities are just complementary dispositions, but we name them differently to keep track of them in our analysis, because it will matter which one we are talking about at any given time.
He stresses that affordances are not relational but dispositional. He then talks about the need for ecological information to solve the problem of us not being able to actually perceive things like the solubility of salt without the need for mental computation.
Ok, so we have the terms. But what can we do with them? Affordances as dispositions - fine. Ecological information - good. But let's resolve some real problems with "salt" for both representational and non-representational psychology.
## Some problems with perceiving 'salt' - from shakers to angry villagers
Salt is salty. Saltiness is a disposition which my effectivity turns into an affordance (if I get the terms right). But what is happening when I'm reaching for some salt to put into my soup?
- Saltiness is nowhere represented in the environment at that moment, yet, it is a part of that process. I'm reaching for it instead of pepper or sugar for a reason. The effectivity of my salt-tasting is a part of me whether I want it or not, but I had to learn that that's what salt does. That is so inherent to my lived experience of salt and saltiness that to think about it with some sort of 'mental representation' - script, schema, plan, frame - is as natural as breathing. But I have a problem. I am not running a structured search through some sort of a repository of frames nor doing a Bayesian probability match as I'm reaching for salt. Salt, as ecological psychology so intriguingly says, is directly meaningful to me. I am solving a task. But without, for lack of a better word, knowledge of salt and its properties, or dispositions, if you wish, it is stripped of all meaningfulness. In this context, renaming a property to a disposition is just doing a job of relabelling an established term without adding much in the way of new ideas. But there are other problems.
- Another disposition of salt is its whiteness. This would be very important to me were I to be asked to pick out salt out of the carpet. But the whiteness of salt is a learned property. These days, there's apparently also Himalayan salt which is pink. When someone tells me to put the salt as I know it into a box labelled 'table salt' and the 'pink salt' into a box labelled 'Himalayan salt' that I never heard of before, something important is happening. When I'm performing those actions, both 'table' and 'Himalayan' become directly meaningful to me but their meaningfulness also includes the new labels, the fact of novelty, etc. And all of that is a part of my experience of them. I think that this again is not mediated by some algorithmic process of matching the perceived objects and conceived concepts to stored representations. It flies in the face of everything we know about the speed of processing and the limits of working memory. But the richness of, for lack of a better word, representation is a part of it. I literally could not accomplish the task without all of that (and much more) being involved. True, I could train a pigeon to do this without this sort of representation but it does not scale.
- Of course, sometimes, I am doing something very similar to running a structured search through some stored representations. For some grains of salt, I may not be sure about whether the level of pinkness justifies it being classed as "Himalayan". So, I may hold it up against the light to examine it, look at it side by side with a more prototypical example of Himalayan or regular grains of salt. I may even ask someone near me to help adjudicate. I experience all of these computationally (that's why the computational metaphor of the mind precedes computers) but the inputs into the 'algorithms' I am perceiving are not computational. I think we can tell an ecological story about this but not without some account of explicit mental operations.
- But I may not even be visually perceiving anything resembling salt. I may just be seeing a white porcelain object on the table with holes in the top of it. Ok, that seems straightforward enough, but I travel between countries a lot and some countries put salt in shakers with one hole and pepper in the shakers with three holes while others swap them around. So, I have to take that into account. Most of the time, I do this without deliberation - direct perception being in operation. But sometimes I pick wrong. I see pepper pouring out of the three holes I want salt from, I see the difference, figure out the problem, remind myself which country I'm in, and go on to pick the right shaker. Of course, this is not something anyone tells you when you cross the border, so I also had to figure out after I made this error a few times. And I have since talked to many people about it as an 'interesting' example of cross-cultural differences. (I am a riot at parties).
So, we have planning, reasoning, learning, error correction and people talking/writing/laughing about all of the above.
And that's even before we get things like metaphor - "I'm pretty salty about that" is a fairly new phrase (at least for me). But I heard that sentence recently for the first time and immediately knew what it meant. How did that instant knowledge come about. There's a new village being built nearby stirring a lot of controversy among the locals. It is going to be called 'Salt Cross' to allude to historical trade routes going through the area. Somebody on Facebook remarked that it's funny because both words describe how the locals feel about it. What's that all about?
## Some theories: Phenomenology, language as a gesture, cognition, and conquest of abundance
When I'm navigating the world, all of that, for lack of a better word, 'conceptualising' is happening, as well. Deliberation, metaphors, arguments are also a part of a directly meaningful world and I am very interested in thinking of my experience of these as not being mediated by complex mixing and matching of abstract symbols. But I'm more likely to find robust treatments in works by a phenomenologist like Merleau-Ponty who talks of the experience of language being similar to brushing off a mosquito off of one's arm. Speaking a language is a sort of gesture on this account and I think there's a lot to recommend that framing.
But if we want to understand the world this gesture is interacting with, we have to take into account the sort of analysis Lakoff did in Women, Fire and Dangerous Things of words like mother or over. Radial categories and idealized cognitive models may not be objects and our brain may not be a sort of computer in which these objects are stored and operated on. But they do describe real experience of the world.
What's more, words, objects, or actions are never on their own. They always combine and recombine. They are differently meaningful to us in different contexts. Perceptually tiny differences have huge conceptual consequences. 'Salt of the Earth' and 'Salt the Earth' differ by but one word, yet, there's a world of a difference between them. I want to have a theory like conceptual integration by Fauconnier and Turner to make sense of these combinations. Ecological psychology can account for these changes when dealing with throwing or running after a football - but when it comes to language, something like constructions and blending seems like a better model. Many people try to formulate these models computationally, but this leaves out most of their richness and malleability. I think the 'ecological' perspective of direct meaning is much more fruitful here. But then ecological psychologists go around insisting that mental representations don't exist.
Unfortunately, the world is not always directly meaningful to us. It presents puzzles and problems. We take wrong paths, trip up, make mistakes of all sorts. When we pause and start to correct them, some sort of a different process kicks into place. Perhaps the most ignored part of Kahneman's book on Thinking Fast and Slow is that there may not actually be any such thing as 'system 1' and 'system 2' but we behave as if there were. We need to acknowledge that this is a part of our world, as well. Any theory of humans that sweeps this under the rug is as incomplete as one that puts it front and centre.
Feyerabend's last unfinished book was called 'Conquest of Abundance' and he described it as follows:
The book is intended to show how specialists and common people reduce the abundance that surrounds and confuses them, and the consequences of their actions.
Ecological psychology's program is attractive because it puts us right in the midst of this abundance as organisms. But it needs to do a better job of addressing problems outside a quite narrow set of disciplines. It is strongest in the fields of perception and motion but has very little to say about what we think of as cognition. There are some hints (https://psychsciencenotes.blogspot.com/2021/02/the-constraints-based-approach-to.html) that this possible but it's not clear how we get from the basics to the higher level stuff.
## The case of digital reading and computer interfaces
I'm currently looking at how people learn to read on digital devices. What is the transition from a printed medium with physical affordances to a digital medium with most affordances mediated by a symbolic interface?
And, of course, we can't just account for the user's behaviour and call the job done. We also have to have a plausible ecological story to tell about the designers of the interfaces (not to speak of the engineers designing the hardware). All of these people (users, designers, trainers) live in a meaningful world surrounded not just by objects and interfaces to be perceived. They are embedded in what I've started calling the "zone of proximal capabilities" (as riff on Vygotsky). They ask friends, show colleagues, follow examples, test theories, make and watch videos on YouTube, read books about good design, etc. An ecological theory that does not take this into account has a very impoverished view of the environment in which the human organism is situated.
I can see how a symbolic account would be appealing. But it fails the mosquito test. All that I described is not mediated, it is as directly available to us, as the action of swatting a mosquito. Once it is learned, that is. And these things have history. Social history. Using the mouse to drag objects on the screen, pinching to zoom on the phone. These are all a part of our direct perception of the environment. But they weren't always. Now interface designers can count on this and that's further changing the environment and its affordances. However, we also need to keep in mind that the ability to navigate that environment is not uniform across the population (just like with the physical environment) and that's why people like me have a job. So that's another thing to solve.
## Explaining this joke about Pavlov is a prerequisite for any theory of psychology
I have a new favourite joke courtesy of Reddit (https://www.reddit.com/r/Jokes/comments/2zvwmj/pavlov_walks_into_a_bar/):
Pavlov is sitting at the bar drinking a beer. Someone walks in and it rings a bell over the door. Pavlov jumps up, slaps his head and cries out: "Shit, I forgot to feed the dogs."
I think the joy of this joke is direct, unmediated by a cognitive process. But it also makes no sense without an account of many layers of richly structured knowledge acquired over time in a social as well as physical environment. And as the comments on Reddit (https://www.reddit.com/r/Jokes/comments/20hilx/ivan_pavlov_is_sitting_at_a_pub_enjoying_a_pint/) indicate, this knowledge cannot be taken for granted. And reconciling that directness of the experience with that sort of rich account of 'knowledge' with all its social consequences is what I want from any theory of psychology claiming to account for a "world as a meaningful place".
## History as weather: A fractal theory of history for Ian Morris, Jared Diamond and CGP Grey
Date: 2021-06-03
URL: https://metaphorhacker.net/2021/06/history-as-weather-a-fractal-theory-of-history-for-ian-morris-jared-diamond-and-cgp-grey/
Categories: Framing, History, Scholarship, Extended writing
Note: This post originally appeared on Medium (https://medium.com/metaphor-hacker/history-as-weather-a-fractal-theory-of-history-for-ian-morris-jared-diamond-and-cgp-grey-45b5503486c5) in 2016. This a very lightly revised version with new formatting for ease of readability (/2021/01/the-nonsense-of-style-academic-writing-should-be-scrupulous-not-stylish/). It preceded the post on historical revisionism (/2019/09/so-you-think-you-have-a-historical-analogy-revisionist-history-and-anthropology-reading-list/) and anthropology of family (/2020/02/its-not-personal-its-family/) but it tackles and elaborates on some of the same themes.
## Outline of the argument
History is often accused of not being sufficiently scientific. To remedy this people try to come up with all sorts of theories of history that try to look like science. They come up with some measurable variable (e.g. availability of domesticable species as Jared Diamond did or energy output as Ian Morris did or size of population, advancement of technology, size of armies, etc. as many others do). This, they hope, explains the past and predicts the future.
Based on the theories, the would-be-scientist-historians build models that seem to make perfect sense but seemingly fall apart when overtaken by events. This makes history seem like it’s not at all scientific. But that’s only because we’re comparing it to the very rare instances in science where long-term perfect prediction is possible. Like the motions of planetary bodies or the behavior of a computer.
But in most cases, science struggles with perfect prediction. In medicine, it’s impossible to predict the exact course a disease will take in any one individual, even though it is often possible to predict what course it will take over a thousand individuals. In biology, it is impossible to predict, which features will be selected over others through natural selection. But it is possible to predict that some will be. But the best analogy, to my mind, is the weather.
Weather is the result of a large number (fractally infinite), perfectly knowable, describable and predictable physical events. But yet we cannot perfectly predict it even in hours or days and not at all in weeks or months. History (or rather any given present in the past or now or in the future) is also the result of a large (fractally infinite) number of relatively knowable, relatively describable and relatively predictable social events. Yet, we cannot predict it very well on any scale worth predicting.
The reason we don’t think of history as science but we think of meteorology as science even though both build models based on observation, known regularities and constants, is because the sort of reliable predictions a historian can make are of no use to anybody, while any kind of even moderately accurate weather prediction is extremely useful to everybody.
Here are some more things about the weather that can be used to view history. On most days, the weather is the same as the day before. You can predict the weather based on the current climate and what you know about the world for a few hours to days in advance (even if you still make errors). It is impossible to predict the exact weather (or history) months in advance but you can predict a range of possible weather patterns (e.g. summers in Europe can be warm or cold but there won’t be snow; political relations in the EU are warm or cold but there won’t be war).
You can also predict trends in changes in the climate but not the exact patterns or consequences or timings of those patterns (e.g. average global temperatures will rise but will it mean it will regularly snow in the summer in France?; disagreements between EU states will continue but will it mean a return to recurring warfare known up to the 1950s?). Most common mistakes in modelling the climate/weather and history come from confusing the local with the global (weather with climate and happenstance with trend) — e.g. during a cold summer, people may question global warming or the fall of the Roman empire was being predicted pretty much throughout its history until it happened hundreds of years later. Both weather and history can be disrupted by freak (Black Swan) events.
What this means is that history can be thought of as much more scientific than is commonly claimed without having to construct models that look like those of climatology.
## History’s attractors
In semi-technical terms, weather is modeled using complexity theory. One of the properties of complex systems is sensitivity to initial conditions. Which means small changes in initial conditions can lead to large swings in the system’s state. This is often described as the ‘butterfly effect’ which is completely misleadingly described as ‘butterfly flapping its wings will cause a hurricane’. [See post on dangers of taking the analogy too far (/2018/11/cats-and-butterflies-2-misunderstood-analogies-in-scientistic-discourse/)]
A much better example might be a weather front moving over a wooded area results in a tiny change of temperature that can then result in a hurricane developing if it was +0.5 degree Celsius or not developing if it was +1 degrees Celsius. Neither state ‘caused’ the hurricane — which was caused by the all the big forces that cause big things like hurricanes. But it was part of an initial configuration that structures the possibilities.
- We have a certain set of contexts in which certain patterns tend to occur and others do not occur at all. With weather this could be climate (local or global) so it doesn’t snow in the tropics or rain in January in the Arctic. In history, people cannot use technologies not yet invented, or generally new agents do not appear out of the blue. [This is the obvious part.]
- Rarely, freak events occur. You may get unseasonably warm weather in the Arctic, hurricane out of season, snow flurry in the tropics. In history, the Mongols, Conquistadors, ISIS came seemingly out of nowhere (and were not in any way predictable by models based on standard assumptions).
- But outside of the freak events, you can predict the range of weather patterns in a given place during a given time period. There’s a range of weather patterns on Earth and weather stays within the limits set by those patterns. And you can do the same for history. Certain types of things are likely to happen (similar to things that have happened) and others are not — e.g. Alien invasions, raptures are unlikely to happen but invasions, civil wars, famines and epidemics are happening all the time.
- It is possible to predict local weather patterns with a decent level of precision based on models in the extreme short term (hours to days) but even these predictions are subject to significant errors due to assumptions made in the models. E.g. it may matter whether a front went over a wooded area or a field. It is possible to accurately predict historical events in the extreme short term — hours to days — simply based on various models of causation (election results, scheduled events, advancing armies) but it is still subject to frequent errors due to assumptions in the models
## Fractal histories
Above, I used the term ‘fractal infinity’. This is not (as far as I could find out) a technical term. It’s a metaphor based on the coastline paradox made famous by Madelbrot’s paper ‘How long is the coast of Britain’. This is the description from Wikipedia (https://en.wikipedia.org/wiki/Coastline_paradox):
the length of the coastline depends on the method used to measure it. Since a landmass has features at all scales, from hundreds of kilometres in size to tiny fractions of a millimetre and below, there is no obvious size of the smallest feature that should be measured around, and hence no single well-defined perimeter to the landmass.
This means that potentially, there can be an infinite (or an indefinetely large) number of measuring units which is the paradox bit. But more interestingly these coastlines have patterns that are similar to each other across scales. This is how Madelbrot (http://science.sciencemag.org/content/156/3775/636) put it:
Geographical curves are so involved in their detail that their lengths are often infinite or, rather, undefinable. However, many are statistically “selfsimilar,” meaning that each portion can be considered a reduced-scale image of the whole.
Why is this relevant to history? I don’t want to overegg the analogy. Mandelbrot’s contribution was a mathematical description of these types of geometric objects and it worked well in certain quantifiable contexts. But that’s not what the study of social objects such as history needs. Those same quantifications won’t work on more analog problems. However, the analogy can explain a huge problem in historical analysis — namely conflation of similar patterns at vastly different levels of magnification.
Imagine history (ie. events as revealed to a human observer over time) as a landscape. As you fly over the landscape from a certain height, the different areas seem very similar or possibly identical (from a large enough distance, Earth will look like a dot). So, from a large enough distance all history will seem like striving for resources of groups of people. So it will make sense describing history as that.
When you zoom in a bit closer, you will see a great inequality in the struggle for resources so you may want to describe history in terms of ability to project power. And you may want to quantify those differences.
But if you zoom in even closer, you will see patterns that look very similar. Even though the groups are of different sizes, they all have very similarly hierarchical organization — with some sort of leadership on top which has to change over time. So, you will try to come up with rules for the formation of these hierarchies.
Yet, as you zoom in even more and you will notice completely different customs, ways of negotiation, ways of legitimizing. So you will postulate a complete incomensurability and simply describe the difference (kind of like a taxonomist biologist).
But then you zoom in even more at the level of individual motivations and you will see things like lust, hunger, aims, struggle for personal success, family — and then you can postulate that we are all the same.
But the problem is that it is actually relatively easy to describe the different levels of magnification in terms of each other because they are sort of like models of one another. An individual’s desires and dreams can be recast in terms of striving for resources; and the behaviors and customs at the level of a kingdom can be talked about in terms of individual desires.
So, the ‘dimensions’ of history depend on the length of the yardstick used to measure it. You can describe the same event from the ‘big man of history’ perspective, the ‘social forces’ perspective, ‘struggle for resources’ perspective or ‘long-duree’, inevitable forces of history perspective. Just like you can describe a chemical event at the level of the structure, molecular interactions or matter transformation.
But all of these levels/perspectives have their own rules and seemingly causal patterns. They also interact with each other but in complex ways that preclude simple reductionism. That is, we cannot say that interactions at one level of magnification cause interactions at another in the same way that an individual brick (or the laying of it) does not cause a wall to exist.
### Aside: Wall metaphor of causality
I’ve come to like the ‘wall metaphor’ of causality. We have completely different intuitions about the causal chain of events resulting in the existence of a wall than we have about the chain of events that come to result in a historical event.
But perhaps rethinking historical events as walls can be helpful — as long as we also keep the differences in mind. We have the brick makers, brick layers, but also the commissioners, approvers, weather conditions, foundations, historical custom of brick making and brick laying — all of those play a role in the sort of wall we’re going to get.
And it is obvious that depending on our perspective the causal chains are going to be completely different. And the same goes for when we try to destroy the wall. We need commensurate forces — big enough force to make a dent, but we also need a lot of small events down to the molecular level. And different perspectives will identify different causal chains. All of them correct but obviously belonging to certain levels. These are obvious (or fairly obvious) when it comes to walls. But not so obvious when it comes to history.
Most good historians actually seem to have the same kind of intuitions about events as most of us do about walls. But they are often seduced by the customs of the genre of history writing into jettisoning their intuitions and coming up with a reductionist perspective instead.
## What this means for the notion of causation in social science
We tend to think of causality in history as happening at the lowest level of magnification and accumulating into the total. And on a most straightforward account, this makes perfect sense. A lot of apples put into a basket + the basket, will make a basket full of apples. But that’s an unhelpful way of looking at causality because it implies a certain level of atomism — a final level of further indivisible measuring units. But that level cannot exist — or if it does exist it has to be so low (subatomic) that the whole cannot be modeled using it. Even if we believe in some sort of terminal particle, it is a fool’s errand (looking at you Stephen Hawking) to try to actually use it to measure object-level phenomena with it.
Which is why the butterfly causing a hurricane is such a poor way of thinking about sensitivity to initial conditions. No one aspect of the physical world ‘causes’ the weather — not the extra +.5 degrees Celsius of temperature, nor the gravitational pull of the moon or the Gulf stream. They constantly interact in a system at all levels. But our ability to model weather is supremely dependent on tiny deviations in measurement.
And the same holds for history and other social systems. No one event causes another event except in the most trivial sense. And an accumulation of little events does not ‘cause’ big events. Because there are no smallest events that could be said to constitute the final level of magnification for initial conditions. But our ability to perceive small events does have huge implications for our ability to model the future patterns of events.
It may have made a difference that Napoleon looked to the left instead of right during a battle thus winning it, which in turn led to his toppling of the Prussian state. Which in turn may have made a difference to the shape of the First World War or the Third Reich. But did that glance left actually cause any of those things or led to those things? Not outside the imaginations of writers of time travel scince fiction. It was just a small difference in the pattern (initial conditions) that made the system come out with a different state at a certain level of magnification.
But on another level, even if Napoleon had never been born, the system may have looked very much the same from a certain level of magnification (European states struggling for resources).
We can have models where big things cause other big things and little things cause other little things. And we must be careful to always know what level of magnification we’re talking about. But we must remember that we are talking about our models of what happens, not what actually happens in its totality. The bigger the models, the bigger errors in our measurement of initial conditions — or rather with the big models we actually have no hope of complete measurement of the initial conditions or any sort of computational tractability even if such measurement were possible.
So even very good and complete models of history have by definition no way of predicting anything with any level of accuracy into any kind of future. And we cannot use their accuracy on past events because our measurement of the data for the past is already filtered by what happened (ie history is written by the winners — or at least, the record keepers). We don’t have that kind of filter for the present or recent history.
We could use what we know from the past to help us sift through the data of the present but that’s where the initial conditions lie that can completely mess up our predictions. Even if we had a lawful model of individual psychology and small event dynamics — in the same way weather modelers have accurate models of molecular and Newtonian object physics — we still could not do a better job at prediction than the weather people can. And in fact, any individual success at prediction is no guarantee of the quality of the model over the long run.
## What this means for Jared Diamond, Ian Morris, Niall Ferguson or CGP Grey
This essay was inspired by a recent podcast discussion (https://www.reddit.com/r/CGPGrey/comments/438ib1/hi_56_guns_germs_and_steel/) of Jared Diamond’s now classic but highly controversial ‘Guns, Germs, and Steel’. But I really started thinking about it when reading Ian Morris’ ‘Why the West Rules: For now’ and Niall Ferguson’s ‘Civilization’. They both suffered from the problem of unconscious magnification refocus.
Ferguson — who’s by far the sloppiest thinker of the three though by no means a worthless one — was the most illustrative example of this. He even could not fix on the idea of the West was for more than about half a chapter.
Morris, on the other hand, was the strictest in his assumptions and measurements — in effect creating two books — the argument and the footnotes. But even he seems to compress time periods and plays around with effect sizes and scales to make the whole thing work. The same thing Steven Pinker is doing in the infuriatingly flawed but worth reading ‘Better Angels of our Nature’.
I think of these as modern historiographical eschatology and it is important to read good anthropologists to fully understand what these people are reading out.
A perfect companion to all of these were David Graeber’s ‘Debt: First 5000 Years’ and most recently for me (although the earliest in publication date) Eric Wolf’s ‘Europe and the People without History’. Also worth reading are the most recent books on Atlantic history. But not of inconsiderable interest is even dross like Pat Buchanan’s ‘Decline of the West’ because it gives an example how many people are thinking about causes.
They all think of a theory of history as a collection of hypotheses about historical causality. But we already know what the causal chains are. Or, the good historians do. But then they forget about most of them when alighting on a good grand theory of history. Jumping from weather to climatology but then using the language of the climate to talk about the weather when they come to the lessons for us today or when they zoom in on the period of history they know really well.
CGP Grey when defending Diamond suggested that people keep criticizing Diamond for all the small things but they never substantively critique his grand narrative. That has also been my initial impression of many of the critiques.
But the criticism of Diamond is actually about him swooping from his heights of resource-utilization level of modeling history down to the level of individual events where he is much shakier and trying to use the same models on the small events. So this leads to a completely inaccurate description of what happened in the Americas — where the guns and steel mattered almost not at all and the germs relatively little (see here for more (/2019/09/so-you-think-you-have-a-historical-analogy-revisionist-history-and-anthropology-reading-list/)).
What mattered in the conquest of the 'new' world were the local politics — the conquistadors became enmeshed in local politics and their victories were results of alliances with other local factions. At the same time, the Portuguese had no chance of anything like that happening on the West coast of Africa or in India (even as they were winning some important naval victories against the Ottomans who had just as much steel, better guns and the same germs).
In some way, Diamond’s critics (who would not be caught dead at an NRA rally) are saying ‘guns don’t kill people, people kill people’. They are saying, real people killed and enslaved other real people — and we must judge what happened for what it happened. It wasn’t the guns, germs or steel that did it. It was our venerable ancestors that did it. And we’re still doing it — albeit in a less mindless genocidal way.
Diamond’s stated intention is to do away with the racial superiority explanation of the difference in the current power arrangements. And he explains a part of it. But he is not careful enough to explore the boundaries of his model. And when applied at the right level, his model does what it sets out to do. But when applied at other levels of magnification, it does exactly the opposite. It provides a way of justifying real bad human behavior as inevitable.
This is perhaps most starkly exemplified in his description of the Rwandan genocide in his other book ‘Collapse’. There he recasts it in terms of simple competition for resources. He may be right. But as recent Timothy Synder’s book on the holocaust argues, that same was true for the Nazi genocide (in the idea of lebensraum). But those were not causes. This involved people like us shooting other people like us. Up close and personal. And other people telling them to do it and benefiting from it in many ways.
Diamond’s models can explain some of the dynamics but they can only do it if they leave a lot of important information out — and the argument is that that’s the information that should matter to us here.
CGP Grey went on to ask, give me a better theory of history. This is a response to that challenge. History is like the weather. This is more a way of thinking about historiography than history itself. Kind of like the language turn in philosophy. We need to be careful with our models, just like we need to be careful with our language. As the saying goes, ‘all models are wrong, but some of them are useful’.
## Stack fallacy, hierarchical structure in science and the meaning of history
As I was writing this, I came across this article on the Stack Fallacy (http://techcrunch.com/2016/01/18/why-big-companies-keep-failing-the-stack-fallacy) which is making the same point in a different domain. You cannot just assume that because you’re good at building all the building blocks, you can build the whole structure. Brick makers are no good at building walls by virtue of being good brick makers and conversely brick layers are not going to be any good at brick making just because they can use bricks to build durable structures. (The post makes the point about databases and CRMs but its even starker at the physical level).
One of the commenters pointed to an old paper by P. W. Anderson ‘More is different’ (https://web2.ph.utexas.edu/~wktse/Welcome_files/More_Is_Different_Phil_Anderson.pdf) that makes a similar point about physics [my emphasis].
The ability to reduce everything to simple fundamental laws does not imply the ability to start from those laws and reconstruct the universe. In fact, the more the elementary particle physicists tell us about the nature of the fundamental laws, the less relevance they seem to have to the very real problems of the rest of science, much less to those of society.
The constructionist hypothesis breaks down when confronted with the twin difficulties of scale and complexities. The behavior of large and complex aggregates of elementary particles, it turns out, is not to be understood in terms of a simple extrapolation of the properties of a few particles. Instead, at each level of complexity entirely new properties appear, and the understanding of the new behaviors requires research which I think is as fundamental in its nature as any other.
A different way of restating this thesis is that the fundamental laws of the lower levels of organization apply at the higher levels (e.g. bricks have strengths and periods of decay and that will impact on how much a wall will withstand or how long it will stand up) but the higher levels have new emergent properties that cannot be easily described in terms of the lower levels. And vice versa, the lower levels cannot be described in terms of the higher levels of organization. However, the structures emergent at all levels can interact with structures emergent at other levels.
What does this mean for history? It can very well matter that a king was a violent drunk who lusted after his subjects’ daughters and wives, and therefore alienated all his followers leading to the collapse of a dynasty. However, when we look at the patterns of falls and rises of things like dynasties, we do not benefit from describing them in terms of individual psychology interacting with their immediate social norms.
That does not mean that the individual psychology and the social norms don’t matter to that individual king or individual dynasty. But falls and rises of dynasties just don’t rest on these issues there are broader patterns and dynamics — kings with highly praised kingly behavior cannot turn around the fall of a dynasty without resources and with external pressures and incompetent despised rulers do not tend to ruin powerful dynasties overnight.
To bring it back to the reimagined ‘butterfly effect’ metaphor. There are valid patterns that can be recognized but patterns are not causes. The exact shapes of the patterns or their varieties are extremely sensitive to initial conditions but the initial conditions are not the causes. Everything that comprises that pattern interacts together (that would include external inputs) to produce all its phases/shapes. It may even be inaccurate to call the pattern itself sensitive to initial conditions. It is our ability to make predictions about its future states and possibly judgments that is extremely sensitive to tiny errors in our observation of what we decide are the initial conditions for the purposes of our prediction making.
With certain weather patterns — like storms — the initial conditions are relatively easy to locate, if relatively difficult to measure. With historical events, it is a little more difficult. It is not always clear where we should look at the initial conditions. It is not uncommon for historians to say things like ‘but the real roots of this crisis lay much further in the past’ — Niall Ferguson’s ‘Civilisation’ is a case study in how badly this can get out of hand.
But historians also tend to freely jump between the levels. So in a recent LSE lecture, Ian Morris talked about long term (in the 1000s of years) trends in war and violence rates. He described an overall long-term equilibrium in violence necessary for the survival of a species. But then he started to talk about our individual natures from an evolutionary perspective, only to turn the talk into a discussion about the motivation of the elites to keep workers happy and alive to produce goods. And then he talked about the British Empire, Cold War, American supremacy and tried to draw conclusions from that.
That’s like saying one of these:
a) ‘there was a storm because it is summer when storms happen’ b) ‘this storm really began when the icebergs melted’ c) ‘there will be a storm on Tuesday, July 3rd, next year because we are experiencing global warming’
Intuitively, none of these statements make sense, even if they may be actually true. Of these, a) is uninformative, b) is also relatively uninformative because we could equally well go back to the formation of the Earth or the Big Bang but other than the arrow of time, we would have little in the way of modelling the causation, and c) is also a statement that is not useful even if it turns out that there actually is a storm on that day. Even if accurate, since our model is consistent with there being a storm on any day in July next year, it is not a useful prediction.
Yet, historians make statements like these all the time — on the surface, they are not this starkly nonsensical but if you peel away the narrative layers and just get down to simple causation, you get things that look like these. At the extreme, you get things like attributing the decline of the Roman empire to lead piping or the rise of the British empire to boiling tea. But even more complex interpretations of past events suffer from similar difficulties.
For instance, post invasion, dismissing all the Bahtist army officers is often cited as the roots of the current violence in the region. But we could plausibly imagine violence of different type but similar scale ‘resulting’ from not dismissing them. Similarly, when things go well in South Africa, everybody’s pointing to the peace and reconciliation for not creating any new resentments and drawing a line under the past. And when they go badly, people wonder if reconciliation was a problem because it did not give people justice or just stirred up trouble.
Ultimately, we cannot say much more than both are things that could happen. We simply have no way of tracking the causal chains at the level in which we could even run the models.
Critics of the social sciences argue that unlike the meteorologists who can build their larger weather models on solid physics, chemistry and geography, historians don’t actually have solid enough psychological models underlying their bigger sociological models. But that is to make the same error. Meteorological models can just about predict the weather tomorrow and the day after while sometimes making huge errors. Historical models are just as good. Equally, historical models can predict larger trends about as accurately as climatology. But we often treat them as if we thought it was possible to predict the weather a year from now.
So what is the point of history then? Its accurate predictions are not very useful and its useful predictions are not very accurate. Any statement like ‘people who forget their history are doomed to repeat it’ are nonsense. Knowing exactly what happened in the past is no better than knowing what the weather was 10 years ago. Other than knowing that anything that happened in the past can happen again, we’re no better off.
But the weather analogy can come to the rescue. On most days, for most people, it does not actually matter what the weather is going to be like tomorrow. Yet, people obsessively check their forecasts. It is interesting. And also it is something to talk about. And that is not nothing. Historical knowledge — you could argue — is even more valuable. It is something to talk about but unlike the weather (unless you think God sends it as reward or punishment), it can be used to help us make sense of today — not in a causal manner but in a narrative one. And the smart historians know that.
## Why I am a feminist: A reading list
Date: 2021-03-08
URL: https://metaphorhacker.net/2021/03/why-i-am-a-feminist-a-reading-list/
Categories: Education, Framing, Gender, Reading Lists, Extended writing
I became a feminist because a woman once told me not to be an idiot and I decided that it was good advice. That was in 1998. But I was all ready to be a feminist long before that, so it really just took a small push to get me over the hump. I was always surrounded by strong women who outshone the men around them, read books as a boy with girls holding their own, later on had women friends who I could respect and like more than most of my male friends.
Yet, I was reticent to apply the label to myself. In those times, it was the women around me who were very sceptical of feminism. So, even when I had doubts about the essentialism of the male/female difference, I was willing to go along without examining the position in too much depth. But looking back, I don't think I was too happy about it.
So it took just a little jolt to show me a new way to reflect on things and I never looked back. Not being a feminist would now feel just as odd as not being a vegetarian. (http://dominiklukes.net/bibliography/masozenyasvobodaduse)
But perhaps I did not always go about it the right way. In the early days, I felt like I had to go back to all my women friends and try to convince them that they really should be feminists, too. Later, I wrote articles supporting political correctness, reviews of Vagina Monologues and other books about gender. And the occasional blog post since. I tried my best.
I have not really felt the need to do any of that recently because it seems to me that the world of today has plenty of voices closer to the action and perhaps I don't have much to add. But then I started looking for a reading list for a friend and was surprised that few of my top choices made the top choices of others' lists.
## The reading list of reading lists
There a plenty of 'feminist reading lists' on offer:
- Harper's Bazaar (https://www.harpersbazaar.com/culture/art-books-music/g19412213/best-feminist-books-every-woman-must-read/)
- The Atlantic (https://www.theatlantic.com/sexes/archive/2013/02/a-reading-list-of-ones-own-10-essential-feminist-books/273337/)
- Oprah (https://www.oprahmag.com/entertainment/books/g25894109/feminist-books-to-read/)
- Book Riot (https://bookriot.com/best-feminist-books/)
- Good House Keeping (https://www.goodhousekeeping.com/life/entertainment/g33577725/best-feminist-books/) (yes, I know!)
They list classics I know or know about like de Beauvoir, Steinem, Friedan, bell hooks, as well as new authors and books that passed me by. They list fiction, manifestos, polemics. Margaret Attwood makes a frequent appearance, some go as far back as Mary Wolstonecraft and the more complete include Judith Butler. There's plenty to read when one goes by the lists.
And I don't really object to any of the items on these lists. Particularly some of the older ones like de Beauvoir or Wolstonecraft are fascinating thinkers even today. But the six I came up with as recommendations appear on none of the lists I looked at. They don't even make the much more comprehensive Wikipedia List of feminist literature (https://en.wikipedia.org/wiki/List_of_feminist_literature). Perhaps I have something to share, after all.
## My list of 6 Books
These are the 6 books that I use to go to for feminist argumentation most often in my mind and the occasional writing. These books are not necessarily radically political but they are radically intellectual. Which, I think, is the main appeal of feminism to me. At its best, feminism is a radical unthinking of the commonplace. And what is more symptomatic of the commonplace than gender? Everything to do with the male/female distinction is steeped in an apparent natural inevitability. And feminism allows us to transcend that inevitability and gives us the inspiration to look for other inevitabilities to be made illusory in its wake.
I tried to find a related video for each book. Here's a whole companion playlist (https://www.youtube.com/playlist?list=PLEl9d2qvKkBvf-bhxWY3gsNpgCvwDSPqZ).
### "Language and woman's place" by Robin Tolmach Lakoff (1972)
"Language uses us as much as we use language" - thus opens Robin Lakoff the world of possibilities with the first line of her book. It it was published in 1972 but it may be best to read it in its 30th anniversary edition with commentary and reflection. This is a foundational book for understanding the depth through which the gender power disparity is reflected in language.
Since then, there have been many books on gender and language from various perspectives - updating and expanding Lakoff's work. Deborah Tannen is one of the more prolific authors worth reading but 'Language and woman's place' is the place to start.
Here's a podcast where Lakoff talks about language and gender (https://youtu.be/E7TKsCd1aRA). Deborah Tannen has written lots of books and has lots of videos. Here's one I like about the differences between male and female speech patterns in friendship discourse (https://www.youtube.com/watch?v=A2xOfpW6xSo).
### "Why so slow" by Virginia Valian (1997)
This is a book that asks the question of what to do when the battle of the minds was won. In 2020, it seems like the battle is starting over again but in the late 1990s it appeared to have been all but over with just a few loose ends to be tied up.
Valian shows that it is the small steps that can add up to significant blockers of progress. The world in which women (at least in certain parts) find themselves today would have seemed far beyond the horizon of the possible as recently as the 1950s. But the chasm seems ever wider and deeper the closer we are to it. 'Why so slow' looks at some of the components of the gap and can help us find levers to perhaps finally close it. In many ways, many of Valian's suggestions have entered the common discourse but it is still worth going back to the source.
My biggest takeaway was that no overt sexism is even required for a society to end up with an imbalance between the sexes. Just small cumulative everyday injustices that may even pass beneath notice of most people involved in them.
Virginia Valian speaks on the topic in a lecture from 2009 (https://youtu.be/FtvV6Bot28Y).
### "Is multiculturalism bad for women?" by Susan Moller Okin and other contributors (1999)
Intersectionality is now on everybody's lips but the history of the fight for equality by various marginalised groups is one of constant tension. For example, women's movements were at different times both deeply intertwined with movement for racial justice and keeping it at arm's length.
This is a book of essays written in reaction to Okin's famous article that underscored the practical difficulty of women's rights when confronted with cultural rights. All contributors start from a concern for women's rights but some even turn the question on its head by asking questions like is feminism good for non-western women? Some question the underlying categories. But none offer easy answers.
When it comes to the utopias hidden behind the placards held by marchers needing to be translated into dirty realities, this is the book to turn to. The title of Martha Nussbaum's essay, "A Plea for Difficulty", should perhaps be a rallying cry for the aftermath of all revolutions just about to tuck in into their children.
I couldn't find any videos about this book but there are good discussions of the topic here (https://youtu.be/yB9baefrHl4) and here (https://youtu.be/sgmvMvrTuC0) even though they do not explicitly reference this debate. I mentioned Nussbaum and here's a great interview with her on gender and development (https://www.youtube.com/watch?v=Qy3YTzYjut4).
### "Marriage: A history" by Stephanie Coontz (2005)
This is not a book about feminism at all. It is a history of the Western side of the institution in which men and women most often used to encounter each other. It is a book that shows how what we know as 'marriage' is determined by the way we talk about it, the surrounding economic realities, and what came just before. Most importantly it shows the constant invention and reinvention of marriage and puts the lie to the 'traditional' notion of marriage invented by 1950s TV.
Stephanie Coontz talks about her earlier book 'The Way We Never Were' in a lecture here (https://youtu.be/MIeAnU7_7TA).
There's also a great interview with her on a New Books Network podcast (https://newbooksnetwork.com/stephenie-coontz-the-way-we-never-were-american-families-and-the-nostalgia-trap-basic-books-2000/). She also wrote a subsequent book about feminism and its reception in the 1960s and is interviewed about it here (https://newbooksnetwork.com/stephanie-coontz-a-strange-stirring-the-feminine-mystique-and-the-american-women-at-the-dawn-of-the-1960s-basic-books-2014-2/).
### "Against Love: A Polemic" by Laura Kipnis (2003)
Kipnis is not against 'love' the feeling but against love as the organising principle of male-female relationships. Where Coontz shows the invention of love as an institutional category, Kipnis takes it apart in its various daily guises that people willingly enter into just to bind themselves to somebody else's nostalgia. Love, when conceived as a social institution, is worse for women than men. Exemplified by the traditional wedding photo of a woman turned subtly towards the male, while the man stares boldly forward. It is not love but the performance of love that can do more harm than good.
I couldn't find Kipnis talking about this book anywhere but here's an interview with her about a later collection of essays (https://youtu.be/1AEySO8Uyug) which shows her insight and subtlety of argument even in the face of a mediocre interviewer.
### "The Gender of the Gift" by Marilyn Strathern (1988)
Not for the faint of heart, this is a dense book of anthropological insight. But it puts the question of women's role in society in a non-industrial perspective. This is both a feminist book but also a meta-feminist book thinking through the dualism that rejecting the dualistic can lead to. It also offers an answer to the question of multiculturalism being bad for women before it was ever asked - the answer being it's complicated.
I'm not sure if reading the first five books would make one ready to read Strathern - to get the wider point one has to read through a lot of technical anthropology. But what is feminism after all if not applied anthropology? Ultimately I think it's worth it - if perhaps a journey best embarked upon with a friend.
Strathern talks about an earlier book on gender and anthropology briefly here (https://youtu.be/afmJIpc7nWk). Here's her later lecture on a different subject (https://youtu.be/xDavTWIKqM8) but showing the anthropological approach to similar issues.
## All other books
All books, when read with gender and sex unthought, can be be part of a feminist reading list. They can be read for what they don't say, as well as what they do say. I wrote about some of my thinking on this here (/2011/04/i-object-a-male-feminists-view-on-the-dutches-of-cambridges-wedding-dress/), here (/2014/05/what-does-it-mean-when-texts-really-mean-something/) and here (/2013/09/storms-in-all-teacups-the-power-and-inequality-in-the-battle-for-science-universality/). And I seem to have written even more than I remember in Czech here: Feminismus a láska.pdf (https://s3-us-west-2.amazonaws.com/secure.notion-static.com/b1a1ca65-125c-4a99-b1be-7b8de5815d3d/Feminismus_a_lska.pdf)
And when all else fails, there's always Buffy the Vampire Slayer (https://www.bufferingthevampireslayer.com/) to turn to.
## The nonsense of style: Academic writing should be scrupulous not stylish
Date: 2021-01-17
URL: https://metaphorhacker.net/2021/01/the-nonsense-of-style-academic-writing-should-be-scrupulous-not-stylish/
Categories: Scholarship, Extended writing, Writing
## The problem with writing advice
The problem with the likes of Steven Pinker and Helen Sword is that they like their own writing way too much. But I don't. Like their writing, that is. [1] I want to get some information from them and I want to get examples and counterexamples for the points they make. I want them to get to the point. I am not reading them for enjoyment, that's what fiction is for. I am reading them to learn what they have to say and I have to wade through a morass of stories, pointless metaphors, geysers of words. They are aiming for eloquence but effluence would be a better term for what their reader gets.
Of course, they get high praise and esteem from their peers, everybody wants to write like Pinker, Gladwell or Sword. And obviously many people buy and read their books. So they must be doing something right. But my claim is that they focus far too much of crafting their sentences and far too little time on crafting their advice.
And what's worse, they promote an environment where 'writing well' with style or panache is seen as a virtue. People are praised for that sort of writing and others are encouraged to emulate them. Sentences like "This was so well written, why can't more academics write like this" proliferate. But it misidentifies the problem. When academic writing is bad, it is not because of the opacity of prose but rather because of the paucity of scrupulous argumentation.
## Defending academic writing
The complaint that academic writing is needlessly dense and abstruse has become a cliche and gets far too little examination. Nobody (except many frustrated readers) is officially complaining that non-fiction writing is needlessly flowery and sprawling across many more pages than necessary. It prides itself on taking the reader on a journey of discovery but it's actually all smoke and mirrors.
There is only one criterion that we should require of academic writing and that is the same requirement we should have of academic thinking. Scruples. Scrupulous writing will let the reader in on the uncertainty of knowledge, and messiness of the process of how knowledge is created. It will not try to write a press-release while presenting its argument. And it will not waste the reader's time by taking them on a journey. Starting every single point it wants to make with a story!
## Starting with academic reading
Academic writing needs to start with the recognition of what academic reading looks like. The vast majority of academic writing is not read like fiction or popular non-fiction. It is read to get the piece of information one needs or to get a gist of what the overall point is. So, having a clear outline with descriptive titles would be much more important than removing unnecessary adverbs. Having an abstract that summarises the key points made in bullet points helps more than using active verbs. Using vocabulary appropriate to the needs of the audience is much more useful than trying to avoid jargon or acronyms.
The first advice you need to give to an academic writer is not to read a book on stylish writing but rather to read how people in their field are writing. Because those are their potential readers. And in their writing, we see what they are expecting. So anything written in the manner they expect will make reading easier for them. Because there is no academic writing as such, there is only writing within disciplines and communities. And barging into a community and trying to change what it's doing without invitation is not stylish, it is rude. [2]
Why are people reading and why are people writing? They are reading to discover the argument that the author is putting forward. And they are writing to put debate the arguments of others and to put forth more of their own. Their success in doing so should be our primary criterion. The success of academic writing is not some abstract readability score but the ability of the peers to identify and debate the argument it puts forward. [3]
## Practical advice on composition: Shorter sentences
This is where we can give some practical advice. And the advice can be summarised in 3 words: 1. Write 2. shorter 3. sentences. Ignore everything about passives, jargon, conversational writing, whatever else the latest guide puts forward. Focus on keeping your average sentence to about 15 words. Not every sentence has to be that short but if you keep the average somewhere between 15 and 20 words, your writing will be easier to read.
Keeping sentences short is not only good for the reader, it is also good for the writer. Often a long sentence means muddled thought. If I find a sentence I wrote that's longer than say 30 words, I often also find that the thought behind it is not clear enough. I had an idea in my mind that had too many assumptions and I had not put enough work into parsing out all the connections.
But not always, sometimes a longer sentence is better for both the writer and the reader. Text that is easy to understand is cohesive as well as coherent and too many short clauses will lack cohesion. The reader needs to know which things link together and making that explicit will make the text better.
The good news is that sentence length is easy to check without having to learn opaque grammatical concepts. We know that people most exercised about passives (http://www.lel.ed.ac.uk/grammar/passives.html) are least likely to actually spot them in the sentence. So why should we expect normal writers to pay attention to parts of speech or fine details of sentence structure. But anybody can paste their text into the free Hemingway Writer (http://www.hemingwayapp.com/) and see if they can shorten the sentences highlighted in red.[4] Doing it regularly will also give you a chance to focus on how sentences are put together.
## Structure is king: Rich outlines
If you really care about your readers, make sure you make the outline of your writing explicit. Don't just mark the normal sections like: Introduction or Conclusions. That is helpful for navigation but not understanding. Put headings inside these sections and have them every few paragraphs to summarise what you're trying to say in those paragraphs. This will aid readers who are not reading your text in sequence to get the gist and also not to miss important points.
You may complain about the youths today not spending enough time on deep reading, but paradoxically, if the structure is marked clearly enough, people are likely to read more of the actual text, as well, because it breaks the work into smaller chunks. Particularly students or novices are often horrified by the walls and walls of undifferentiated texts they are asked to climb. Dividing it into smaller, labelled chunks makes the text more approachable and more likely to be read.
This rich outlining can be done while planning your writing, during writing or during revision. I often don't start with an outline because I'm not yet clear on exactly what I want to say. But I always create an outline at some point. It helps me discover all the missing and inconsistent bits that are so easy to miss in an extended narrative.
And, finally, if you have an outline, share it with the reader. Put it at the top of the text. It could be a table of contents or just a list of bullets. Better still, put it in the abstract. It is infuriating, reading an abstract that says what the paper tries to show and not what it actually showed. [4]
## All you need
And that's it, clear explicit structure and moderately shorter sentences. No stories, no metaphors, no flourishes. No avoidance of passives or reduction of adverbs. No worries about technical language. Just these two. They will not only make the academic writing easier to read, they will also make it more scrupulous.
## Self-criticism
Do I write as I say? Not always and not always perfectly. Some of my sentences do get long and my composition tends to the flowery. For example, the sentence about the importance of shorter sentences is the longest one of this post at 43 words. I thought about breaking it up but I liked its rhythm, so I left it. But my average sentence is 14 words and this makes the overall readability higher.
But that is the lot of anyone who writes about writing. The more strident your strictures, the more likely you are to fall afoul of them (https://www.polysyllabic.com/?q=node/168).
## Notes
[1] For those who do not routinely read books on writing. The title of the post refers to a book by Steven Pinker "The Sense of Style" which is a lot to wade through for not very much useful advice. The other book I complain about is Helen Sword's "Stylish Academic Writing", which overall has better and more actionable advice than Pinker. But the problem with Sword is that many of the examples she sets out to change actually don't need changing that much. And the examples she singles out for praise are questionable.
[2] The advice about how important it is pay attention about the language of the community you're writing for is best summarised by Larry McEnerney in his talk about Writing beyond the academy (https://www.youtube.com/watch?v=aFwVf5a3pZM&list=PLEl9d2qvKkBslX4kb0bG2mWmiSKYDGBdC&index=3).
[3] The importance of thinking about your writing as communicating with interested peers is often made by Thomas Basbøll) on his blog Inframethodology (https://blog.cbs.dk/inframethodology/). But it is also the core of the argument of this paper by Cathy Birkenstein defending Judith Butler (https://www.jstor.org/stable/25653028?seq=1) against the charge of incomprehensibility.
[4] The Hemingway App (http://www.hemingwayapp.com/) also tries to identify passives and adverbs. Feel free to entirely ignore its advice. Also, it computes a readability score which is useful. But all it does is count letters in words and words in sentences. It is easy to game just by randomly inserting periods through out the text. ****
## Metaphors and freedom: On Tolkien's notion of allegory vs applicability
Date: 2021-01-01
URL: https://metaphorhacker.net/2021/01/metaphors-and-freedom-on-tolkiens-notion-of-allegory-vs-applicability/
Categories: Framing, Metaphor, Extended writing
Tags: Framing, Metaphor
On rereading Tolkien's Lord of the Rings, I was struck by this passage in his foreword to the second edition:
I cordially dislike allegory in all its manifestations, and always have done so since I grew old and wary enough to detect its presence. I much prefer history, true or feigned, with its varied applicability to the thought and experience of readers. I think that many confuse 'applicability' with 'allegory'; but the one resides in the freedom of the reader, and the other in the purposed domination of the author. http://ae-lib.org.ua/texts-c/tolkien__the_lord_of_the_rings_1__en.htm (http://ae-lib.org.ua/texts-c/tolkien__the_lord_of_the_rings_1__en.htm#00)
Tolkien is reflecting on the many reactions to his work and people trying to situate the happenings in the Lord of the Rings into the context of the Second World War or later the Cold War. He claims that he had no such intentions and takes some time to analyze what would have had to have happened in the books if he had wanted to make that sort of point.
But the idea of history opening up more opportunities for readers than allegory is a useful one. Allegory is just a complaint, an empty gesture aiming to show something in a purportedly "truer" light than a mere description. But in fact, it aims to constrain the lesson that can be learned, rather than expanding it. This neatly underscores the central thesis of this blog that metaphors and the conceptual frames they rely on need to be negotiated. Taken apart, examined and put back together.
Tolkien's complaint helps me express much more succinctly my own dislike for allegories and dystopias such as Orwell's Animal Farm or 1984. There is so much more to learn by studying the complex case studies of history and the thick descriptions of ethnography - which in the best cases reside in the same work. What did Orwell add to what was already known? Nothing. He only took away. He took away the complexity of the situation and he absolved his readers of tackling the more complex issues involved.
But luckily for us, we are not at the mercy of other people's imagination. My thesis is not just normative. I do not just say that we should always examine our metaphors and negotiate their mappings - even though it is a recommendation I do make. This negotiation is what we actually already do as part of the process understanding metaphors. And we don't just do it at the moment of reading a particular metaphor. We do it over the course of our life. We think, rethink, we have conversation with others, we are exposed to competing accounts - even in seemingly totalitarian contexts.
In fact, allegory is itself an example of this process. Because what is allegory if not an extended metaphor, negotiating one view of the world. It is a narrative metaphor, just like satire but without the humor. Its rhetorical purpose is one of persuading us to the author's view of the world. And it may be slightly more underhanded than a mere historical description by constraining the avenues of applicability, to use Tolkien's term. But ultimately, even a historical description or an ethnographic account use the same processes of framing and reframing. Describing the world from a perspective of the author's context.
That's why each generation writes new histories, even if the actual facts do not need revision (though they often do), the new perspective of the time demands a retelling. All of a sudden, there are more women in history, more people of color, history is not happening just to old white men. And how did this come about? Did we all of a sudden notice things we missed the first time around and changed our view of the present? No, we just find new things worth talking about. This in turn does lead to new discoveries of fact, but these will then be reassessed as the narrative demands of the day change.
Metaphor, in this sense, is not a figure of speech, it is a process that expresses our constant reengagement with the world. Metaphors do not give or take away our freedoms of thought or action. They are the means through which we express these freedoms. All the way through. Across books and debates.
Hobbes spent a whole section in the Leviathan complaining about metaphors and their power of misdirection. Without ever once acknowledging that the book itself is one giant metaphor or that he is using the metaphors of ignes fatui (will o'the wisp) to do his hatchet job on them. But he was not wrong, metaphors can be used to cover up parts of what we express as well as to shed new light on it.
But they never have the last word. Because we never stop talking, writing, engaging. Finding new ways of projecting concepts onto each other, coloring our perceptions with lenses made from one framing or another. And this includes the very process of metaphor use. Hobbes or Tolkien are just one of many weighing in on the use of metaphor. Some say that metaphors are just froth on top of good literal truth, others that metaphors are the only way to true understanding.
But metaphors are always the journey never the destination. Even if we sometimes feel their power so strongly that we cannot for the moment imagine anything else, there's another metaphor or allegory just around the corner, just as seductive and completely contradictory. And then a third taking parts from the other two. And it is through this constant reassessment and retuning of our understanding as individuals and as groups, that we live our lives. As free or as tied down as we can be. Metaphors are there to help us along, but we are the ones who wield them. Constrained by the limits of metaphor, having to negotiate the journey with them and around them, sometimes stuck as if in treacle but never chained to them for good.
## No back row, no corridor: Metaphors for online teaching and learning
Date: 2020-06-28
URL: https://metaphorhacker.net/2020/06/no-back-row-no-corridor-metaphors-for-online-teaching-and-learning/
Categories: Education, Extended writing
Tags: featured
## Publication note
An earlier version of this was published in the Oxford Magazine (https://staff.admin.ox.ac.uk/oxford-magazine) No 422. This post expands certain sections based on questions and feedback I received following the first publication of the piece. It is also available on Medium (https://medium.com/metaphor-hacker/no-back-row-no-corridor-metaphors-for-online-teaching-and-learning-9628f164fc37?source=friends_link&sk=27f6f0666a88d9951c5d16a28d0e00e7).
## The state of digital dislocation
The current state of digital dislocation is forcing us to reevaluate what is the essence of teaching and learning. The “grammar of schooling” [1] has been taken away from us and we are forced to learn a new dialect by immersion with just a few phrasebooks, hastily pulled off the shelf, to guide us. Digital teaching is still teaching but it is teaching with an accent, one where we’re still trying to acquire enough fluency and idiomaticity to feel completely at home. When we add to it the culture shock of being in a new situation without any of the familiar cues, sights, sounds and smells of our native environment, it is not surprising that many people are feeling stressed and long for a swift return to “normal”. But it is also no surprise that many others are examining the current situation and finding the new land to be one of endless opportunity and thinking of establishing a permanent residence or at least buying a holiday home.
At one extreme, we are hearing voices calling online learning “clearly inferior,” lacking the essential personal contact that defines the University experience and asking whether the cost, expressed in fees, is too high. At the other pole, we hear “online teaching is clearly better,” doing away with all the distractions and deadweight of spaces, commutes and providing the focus so essential to learning. The same person can find themselves taking either position depending on the stage of culture shock they are living through at the moment. Both of these perspectives were reflected in an eloquent summary by Ray Williamson from the Oxford Student Union in a recent issue of the Oxford Magazine.[2] Here, I’d like to elaborate on what is at stake and look at ways of conceptualising the different perspectives.
## Making sense of digital with affordance metaphors
I suggest that the two divergent views can best be reconciled when we contrast the affordances of the physical and virtual environments in which teaching and learning take place. By affordances I mean those features of the environment that present themselves to us for direct action and interaction and thus make the world around us meaningful and define what it means to live in the space we’re in. Affordance is a concept fundamental to design thinking and interaction and ignoring them is the most frequent cause of failure both in digital and physical products. [3]
The best way I found to bring the contrast between the physical and the virtual into focus are two metaphors that can be summarised as “No back row” and “No corridor”.
## “No back row”
“No back row” expresses mostly the potential of the online experience to be positive for learning: the digital space is the great equaliser, no student is left hiding in the dark corner of the room, everybody’s contribution is coming from the front. This leads to higher engagement with the study material, and better learning. It is so powerful that the American online course provider 2U trademarked the slogan as part of their corporate philosophy [4]. Of course, just because it has the potential to be beneficial for learning, it doesn’t mean that we can just put the same course online and get its benefits. We have to design the online courses to take advantage of this. Nor should we be mislead by the visual metaphor of the Zoom call that 2U use on their corporate page. This applies to an entirely forum-based course, as well. The very fact that they have to engage with the content may put additional demands and stresses on students that will require support. This is in addition to the issues that are captured by the ‘no corridor’ metaphor.
## “No corridor”
“No corridor” reflects the largely negative aspects of the virtual when contrasted with face to face. It reflects the lack of physical and social space connecting the learning situations. There are no natural landmarks to guide us, no flow of the crowd to follow. Everything has to be scheduled, bookmarked or emailed. There is little serendipity and no feeling of just “being there”. This makes it easy for a student to disappear and find themselves in “no row” at all. The physical space is doing a lot of work that is beyond the conscious notice of educators and programme administrators and allows them to be less specific in their instructions and leave things to ‘work themselves’ out without realising it. Their planning may be meticulous and painstaking but it is always framed by what the space affords them when it is filled with students, signs, and other signals that may feel almost subliminal. This can be easily seen when we compare the instructions students receive before arrival (what to bring, where to come, what to expect) and when they arrive which may be as little as a time table followed by ad-hoc announcements. And this comparison may gives clues to some aspects of mitigating the downsides of ‘no corridor’.
## Affordances of the physical vs virtual
Photo by Lucrezia Carnelos (https://unsplash.com/@ciabattespugnose?utm_source=medium&utm_medium=referral) on Unsplash (https://unsplash.com/?utm_source=medium&utm_medium=referral)
Luckily, we can mitigate the downsides of the virtual and amplify its benefits, if we pay careful attention to the affordances of the physical. There are successful ways of making up for the lack of the corridor’s hidden contribution to the learning process but we must avoid taking the normal environment in which learning takes place for granted. We rightly focus on personal relationships as essential to learning but as we saw above it is easy to underestimate the power of the spaces in which they are situated.
In the physical space, it is much easier to just follow the flow of the environment and learn, without realising it, by reflecting others’ reactions to it. There are spaces laid out so obviously that our use of them passes completely beneath any level of conscious notice. We do not need to deliberate on how to open doors, sit facing the speaker, not to sit in a seat already occupied. And where there are issues (locked doors, missing markers, drilling outside the window), we have established scripts for coping and frames for interpreting them.
None of these features are present in the virtual environment. Every action (at least initially), requires the effort of directed attention. We need to learn the “interfaces” of Zoom, establish routines of where to ‘find the link to join’, keep track of bookmarks for the learning materials, and manage actual time for virtual events and assignment deadlines. All of this virtual effort is taking place in an actual physical environment where we are the only person engaging in the activity. What’s worse, when we study or teach virtually, we do not appear to the world around us any different from when we idly browse the web or are binge-watching a TV show. We then have to negotiate with that environment and people in it in ways that travelling to ‘school’ or the ‘library’ does for us without any words having to be exchanged other than ‘I’m going to class’.
It’s no wonder many are finding themselves more stressed, tired and downright disoriented. But equally, to no one’s surprise, there are many who are thriving without the extra burden of the physical space which they may have found too overwhelming, full of distractions and uncertainties. We know that not all students cherish the demands of the physical spaces into which attending a university thrusts them; those who only feel comfortable huddled in the back row or for whom passage through the corridor is an exercise fraught with anxiety. Universities have ample built-in support structures and processes (albeit imperfect) for the latter but none for the former.
## For a successful online learning experience
Yet, we know that it is possible to build a sense of “being there” in fully virtual environments and it is also possible to establish durable personal support relationships. This was possible even before the rise of Zoom or Teams as the success of Open University can attest but now it is even more within reach. Perhaps the most powerful indications of this are coming from the successes of telemedicine and even online psychotherapy. Many patients are finding that their one-on-one experience with a therapist is enhanced without the stressful overhead of travel, sitting in waiting rooms, walking through crowds, etc. [6]
Telemedicine also shows the way when we think about the heterogeneity of needs and inclinations. It is clearly not always appropriate to conduct therapeutic interventions over Skype but it is sufficient or even superior in more instances than may have been thought before the current situation made them a necessity. Do we think that education is radically different, here?
What does a University have to do to make the most of the benefits of ‘no back row’ and minimise the downsides of ‘no corridor’? What does the individual educator? The solutions are surprisingly simple and non-technical. Above all, we need to realise how much we can leave unsaid because the physical environment says it for us and then make it explicit in the virtual setting. We need to communicate more clearly and more frequently. We need to design the virtual learning spaces to minimize unnecessary cognitive load, structure information better, pay attention to navigation and consistency. We need to constantly fine-tune the balance between information overload and not enough information. We need to build structures that support the students who are struggling with the technological as well as personal aspects of learning.
## New roles for the relationship business
Photo by Brooke Cagle (https://unsplash.com/@brookecagle?utm_source=medium&utm_medium=referral) on Unsplash (https://unsplash.com/?utm_source=medium&utm_medium=referral)
But ultimately and most importantly, we need to realise that educational institutions are not in the information business, they are in the relationship business (to borrow a metaphor from the media critic Jeff Jarvis [5]). It is easy to deploy an army of learning technologists and media production specialists, and think we’ve done virtual teaching justice. But online teaching requires other support roles and activities than just those leading to the deployment of “tech”.
There need to be roles whose main job it is to make sure students are opening the right virtual doors and sitting facing the right way in the virtual learning spaces. There need to be roles that pay attention to the real physical spaces and social situations on the other side of the Zoom call. When students are on campus, so much of this is done for us by the affordances of the space built up over centuries and so ingrained into our conceptual and perceptual systems that interacting with them feels to be a matter of instinct.
When all we have is emails, forum posts, webcams and the screen, we need to put in additional work to make up for this. Over time, it will come to seem as natural as what we have now but not without the initial effort. For instance, it is not anyone’s job to explicitly make sure students socialise with others in the physical environment. We don’t ask students if they “went out for a drink” with others when they’re on campus, but perhaps, it needs to be somebody’s job in the virtual situation. [7]
## Sources of learning
Luckily, we have ample models of successful practice to draw on. The Open University is one such, Oxford’s own Continuing Education department is another. Private online education providers such as GetSmarter / 2U, who provided the first part of the metaphor, are others.
As far back as 2009 before Zoom or video conferencing, I taught a module on language and education in a physical setting followed a year later by a similar module in a fully online course for teachers. I was struck, when reading the final essays, how much more the online students seem to have engaged with the subject.
In the physical space, I had a feeling of engagement during my seminars with the students. But the ‘feedback’ I was getting from them hid the relative shallowness and unevenness of their engagement. I never saw the online students in person, so I had to design the course to get this feedback in other ways and I could easily see where all individual students were and guide them back in the right direction if they seemed to be floundering. It was more work for me and them but the learning gains were there to see.
The lessons of this anecdote are supported by research evidence and by experiences of educators the world over [8]. We do not need to provide inferior experiences to students just because they are not in the same room as us.
Eventually the world of university teaching and learning will return to “normal” but we should be mindful that culture shock happens on returning home, as well.[9] We can take advantage of what we learned during this forced sojourn in digital lands to develop a more robust bi-cultural approach to teaching by blending the best of both worlds.
## Footnotes
[1] Tyack, D.B. and Cuban, L., 1995. Tinkering toward utopia: a century of public school reform. Harvard University Press, Cambridge, Mass ; London.
[2] Williams, R. 2020. “Students and remote learning” Oxford Magazine, 421, Trinity.
[3] Norman, D.A., 2013. The design of everyday things. Basic books, New York, N.Y.
[4] No Back Row | 2U [WWW Document], n.d. URL https://cdn2.2u.com/about/no-back-row/ (https://cdn2.2u.com/about/no-back-row/) (accessed 6.8.20).
[5] Jarvis, J., 2012. What the media can learn from Facebook. The Guardian, 15 February 2012, sec. Media Network. https://www.theguardian.com/media-network/media-network-blog/2012/feb/15/what-media-learn-facebook (https://www.theguardian.com/media-network/media-network-blog/2012/feb/15/what-media-learn-facebook).
[6] These two recent pieces summarise the pros and cons of mental and physical health interventions and point to relevant research.
Joyce, N., 2020. Online therapy having its moment, bringing insights on how to expand mental health services going forward [WWW Document]. The Conversation. URL http://theconversation.com/online-therapy-having-its-moment-bringing-insights-on-how-to-expand-mental-health-services-going-forward-136374 (http://theconversation.com/online-therapy-having-its-moment-bringing-insights-on-how-to-expand-mental-health-services-going-forward-136374) (accessed 6.8.20).
Novella, S. 2020. It’s Time for Telehealth. NeuroLogica Blog. URL https://theness.com/neurologicablog/index.php/its-time-for-telehealth/ (https://theness.com/neurologicablog/index.php/its-time-for-telehealth/) (accessed 6.8.20).
[7] Redmond, P., Heffernan, A., Abawi, L., Brown, A., Henderson, R., 2018. An Online Engagement Framework for Higher Education. Online Learning 22. https://doi.org/10.24059/olj.v22i1.1175 (https://doi.org/10.24059/olj.v22i1.1175)
[8] The following systematic reviews show that online higher education is at least as effective as offline education when it comes to learning outcomes.
Means, B., Toyama, Y., Murphy, R., Bakia, M., Jones, K., 2009. Evaluation of Evidence-Based Practices in Online Learning: A Meta-Analysis and Review of Online Learning Studies, US Department of Education. US Department of Education.
Nguyen, T., 2015. The Effectiveness of Online Learning: Beyond No Significant Difference and Future Horizons 11, 11.
Pei, L., Wu, H., 2019. Does online learning work better than offline learning in undergraduate medical education? A systematic review and meta-analysis. Med Educ Online 24. https://doi.org/10.1080/10872981.2019.1666538 (https://doi.org/10.1080/10872981.2019.1666538)
[9] Gaw, K.F., 2000. Reverse culture shock in students returning from overseas. International Journal of Intercultural Relations 24, 83–104. https://doi.org/10.1016/S0147-1767(99)00024-3 (https://doi.org/10.1016/S0147-1767(99)00024-3)
## It’s not personal, it’s family: Kin, strangers, guests, and the complexity of social obligation
Date: 2020-02-16
URL: https://metaphorhacker.net/2020/02/its-not-personal-its-family/
Categories: Anthropology, History, Extended writing
Tags: featured
## Brooks on the alternatives to nuclear family
Tyler Cowen (https:marginalrevolution.com/marginalrevolution/2020/02/the-importance-of-family-structure-was-the-nuclear-family-a-mistake.html) called the extended essay by David Brooks called ‘The nuclear family was a mistake’ a “so far the best essay of the year with many fine and subtle points”. And he’s not wrong. Brooks who has frequently been caught embellishing data to make a point (https:%3C/span%3Estatmodeling.stat.columbia.edu/2015/06/16/the-david-brooks-files-how-many-uncorrected-mistakes-does-it-take-to-be-discredited/) does a very good job with his sources. And most importantly he transcends his natural middle-of-the-road conservative leanings to modify his views to fit the data rather than the other way around.
His portrayal of the current social environment is well worth reading and rereading:
“If you want to summarize the changes in family structure over the past century, the truest thing to say is this: We’ve made life freer for individuals and more unstable for families. We’ve made life better for adults but worse for children. We’ve moved from big, interconnected, and extended families, which helped protect the most vulnerable people in society from the shocks of life, to smaller, detached nuclear families (a married couple and their children), which give the most privileged people in society room to maximize their talents and expand their options.”
But he avoids and indeed criticises the conservative instinct to demand a return to the traditional family:
“Social conservatives insist that we can bring the nuclear family back. But the conditions that made for stable nuclear families in the 1950s are never returning. Conservatives have nothing to say to the kid whose dad has split, whose mom has had three other kids with different dads; “go live in a nuclear family” is really not relevant advice.”
And he also points out the deficiencies in the liberal response:
“Progressives, meanwhile, still talk like self-expressive individualists of the 1970s: People should have the freedom to pick whatever family form works for them. And, of course, they should. But many of the new family forms do not work well for most people—and while progressive elites say that all family structures are fine, their own behavior suggests that they believe otherwise.”
His summary feels very apt to the situation (although it is important to note, that there are many progressive thinkers who are much closer to his ideas than he admits):
“while social conservatives have a philosophy of family life they can’t operationalize, because it no longer is relevant, progressives have no philosophy of family life at all, because they don’t want to seem judgmental”
Brooks also very aptly formulates the current state as a paradox:
“Our culture is oddly stuck. We want stability and rootedness, but also mobility, dynamic capitalism, and the liberty to adopt the lifestyle we choose. We want close families, but not the legal, cultural, and sociological constraints that made them possible.”
In fact, the solution he proposes, created, forged families - a redefinition of kin, is more socially progressive than conservative.
“This is a significant opportunity, a chance to thicken and broaden family relationships, a chance to allow more adults and children to live and grow under the loving gaze of a dozen pairs of eyes, and be caught, when they fall, by a dozen pairs of arms. For decades we have been eating at smaller and smaller tables, with fewer and fewer kin. It’s time to find ways to bring back the big tables.”
This is an essentially progressive vision tinged with a fair bit of the conservative communitarian nostalgia. It is the same nostalgia for the imaginary of togetherness that drove ‘Bowling alone (https://en.wikipedia.org/wiki/Bowling_Alone)’ to such popularity 20 years ago. It is neither venal, moralistic, nor unrealistic. Brooks is merely describing what exists and has always existed in one way or another. He then takes a turn reminiscent of Margaret Mead and says, let’s make that the new normal. And he seems to find the sweet spot that has the potential of becoming a meeting point for the utopian imaginaries of both conservatives and progressives.
What Brooks leaves out are the limits of these family-like structures when it comes to dealing with those outside them. And unsurprisingly he also fails to mention the possible role the state can and perhaps must play in tying them together.
Brooks does an admirable job of engaging with the anthropological literature. He does not just insert the obligatory James C Scott reference so beloved of certain kind of libertarian thinker, he reads more widely and more deeply. But, as always, there’s more. Here I’d like to bring in some more anthropological perspectives to enrich and somewhat complicate Brooks’ vision. This is not to negate what he says or dismiss it as erroneous. All I’m trying to do is expand the picture slightly.
## Kin, guests and strangers: From baseline communism to complex webs of social obligation
The one anthropologist he does not mention is David Graeber (https://en.wikipedia.org/wiki/Debt:_The_First_5000_Years). Graeber called the kind of mutuality Brooks is after ‘baseline communism’. When communism is defined as ‘to each according to their needs, from each according to their abilities’ it is often decried as unworkable because it ignores human proclivities for cheating. But, in fact, as Graeber points out, it describes perfectly one very familiar environment: the family. Communism is not a question of property, it’s a question of obligation from one to the many and the many to the one. This always exists on some sort of spectrum. As Graeber describes, there’s never complete abandonment of private property (even in the most egalitarian societies, there are some things people can call their own), nor complete abandonment of supporting people’s needs (even hard-nosed captains of industry will give each other breaks under certain conditions).
This obligation is unconditional but it is also constrained. To help us understand the constraints, we could simplify the sources of obligation by dividing the social world into three classes: ‘kin’, ‘guests’, and ‘strangers’. Kin and strangers are mostly stable categories whereas ‘guests’ are inherently dynamic and transient. ‘Guest’ is a stranger who becomes temporary ‘kin’ in the sense that the obligations for ensuring the wellbeing of ‘kin’ transfer to them on a limited basis but often in a ‘lavish’ manner. One of the inventions of the ‘post-industrial world’ is our ability to deal with ‘strangers’ without having to confer kin-like privileges on them. The ability to associate and collaborate with strangers who are not your guests is what made industrial capitalism as we know it possible.
But it is also what makes socialism possible. Once the complex webs of mutual support through extended kin networks have been torn up, the state can step in and substitute for this obligation. When many progressives talk about the duty of care of the state (or smaller collectives), they essentially claim this sort of familial role for the state. Conservatives, on the other hand, view the state more as a meeting ground for groups of strangers. For them, the state as such should have no power to treat anyone as guests.
This tension is present even in socialist-leaning countries such as Sweden or Germany which are willing to provide for their citizens - who are otherwise strangers to each other - some of the kind of support traditionally reserved for kin. They are also willing to provide hospitality to new arrivals but very much as guests. What they are struggling with is the process of conferring the kin-status on these guests.
Brooks mentions how many captives of the Native American nations in the 1700s did not want to return back to the ‘civilised’ world from which they were taken. But he omits to mention that these groups often had elaborate ways of transferring strangers from captives to guests to kin. Be it through marriage or adoption, even enslaved war captives could (sometimes enmasse) made into ‘one of us’. Japan, for instance, still widely practices ‘adult adoption’ which was also very common in Ancient Rome and is one of the ways of achieving this that is not available to the state.
Brooks talks about what was lost and he is not wrong. He is also not blind to the fact that the support the kin networks provided was often opperessive and frequently rested disproportionately on women. But he still perceives it from the perspective of a homestead - he talks about the individual groups as if they floated in a vaccuum.
This is where Brooks would have benefited from engaing with another precursor of his thinking Karl Polanyi. Polanyi critiqued our view of the industrial revolution as merely a matter of technology plus capital. He wrote his magnum opus ‘The Great Transformation’ over 80 years ago - long before the 1950s ushered in the nuclear family revolution Brooks blames on current ills. Polanyi traces the problem much further back to the needs of early industrial capitalism. And he also relies on the ethnography of his day to contrast the ‘mutual support’ of traditional societies with the manufactured individualism of the industrial age.
## The emotional pull of the ‘mutuality of being’: For good or for ill
Polanyi does not dwell on the emotional aspects of the material support networks but the nostalgia is clearly there. Brooks, on the other hand, can’t get over the emotional impact of personally experiencing being a member of the kind of group of mutual support that he proposes. And he is not wrong to point out the strong emotional pull of the mutuality of the neighbourhood. Nobody does it better than 2PAC when he sings about the feelings of returning to his old problematic neighborhood in ‘My Block’.
My neighborhood ain’t the same
Cause all these little babies goin crazy and they sufferin in the game
And I swear it’s like a trap
But I ain’t given up on the hood it’s all good when I go back
Hoes show me love, niggaz give me props
Forever hop cause it don’t stop… on my block
and talking about what it’s like being away:
In my heart, I felt alone out here on my own
I close my eyes and picture home… on my block
Brooks quotes ethnographers such as Marshall Sahlins and Monica Wilson on the ineffable nature of the connection within kin groups. People in them experience “a mutuality of being” (Sahlins) and are almost ‘mystically dependent’ (Wilson) on one another. This is easy to overlook in more institutionally focused accounts of ‘kinship’, so Brooks is right to emphasize it but it’s not all there is.
The emotional support mutuality provides is well known and is present even in situations where it is harmful to the individuals. In describing a materially and, by his account, mentally and socially deprived society in a remote Apalchian community, Robert Edgerton, reports:
Despite the absence of any kind of ritual, ceremony, or community-wide activities, these people were fiercely loyal to their hollow and their way of life. Even those few who could emigrate, like a boy Gazaway took away from the hollow for a brief period of schooling, preferred to remain in Duddie’s Branch. They could also express great love for members of their families, and even for an outsider like Gazaway. They had pride, dignity, courage, and generosity.
The affective power of the nearly mystical ‘mutuality of being’ can exert a strong attractive force on a reader who is not enmeshed in such strong ties. This makes it easy to forget, that tightly-knit communities are not always idylic:
“some small-scale populations do not effectively solve the problems they face, and sometimes the very culture that should sustain them and enhance their well-being instead produces fear, apathy, isolation, and degradation.” from Sick Societies by Robert Edgerton
Edgerton also has an agenda of his own but his account is an important antidote to the opiate of anthropological utopia.
My main point is that while the emotional impact of mutuality is substantial, it alone is not enough to account for all the elements that we see in the tripartite ‘kin/guest/strager’ distinction. And neither are social norms as traditionally conceived; viz norms being the combination of unwritten rules, explicit laws, and various forms of enforcement. It is instead a cognitive perception of how the world is. It’s not that people fulfil obligations because of fear of sanctions. It’s because they cannot imagine not doing so. It is simply against a very basic fabric of their being.
## It’s not personal, it’s family: Ties that moor us and bind us
I spent six month working on projects in Timor-Leste, one of the poorest countries in the world, where people frequently don’t have enough food during certain times of the year. Yet, there is almost no homelessness and festivals are common even in the poor areas. How can this be? The answer is extended family. The Timorese large family networks provide material and emotional support for all their members.
But this also makes working on projects quite difficult. The Timorese are no less intelligent or competent than any other people I’ve worked with. But their family always comes first. We’re not talking just about sick children or bereavements but also festivals and other family gatherings - which are not infrequent. Combined with what often seems like a sudden appearance of these events, running projects when a key staff member can disappear at a moment’s notice is often a frustrating experience.
And it’s not even that the person wants ‘go to a fun party’ instead of doing ‘dull work’. Often attending such events can be both emotionally and materially draining. Resisting the pull of the social obligation is like resisting gravity. Sometimes it keeps us grounded, sometimes it throws us flat on our face. To help my non-Timorese colleagues conceptualise this better, I came up with the mantra: “it’s not personal, it’s family”. In the same way that the American “it’s not personal, it’s business” is used to explain or excuse behavior against the norms of sociability, so can the Timorese “it’s not personal, it’s family” be helpful to understand the sort of behavior that the individualistic mindset perceives as a breach of contract.
This, of course, is not unique to Timor-Leste, nor is it unknown in the cross-cultural literature. These family networks can also be harnessed to economic benefit in the capitalist system. Many of the successes of the East and South East Asian diasporas around the world are due in large part to their ability to harness the mutuality of support across long distances. A small enterprise that can rely on the virtually free labor and trustworthy sources of credit or supplies provided by kin has a larger chance of success, particularly is the boundary between the work and personal life is very flexible. And the kin can expect mutual support back - extending across the globe through remittances and other forms of support. Not based on a simply return-on-investment calculus but on bonds unconditional mutuality.
But these same networks often do not work as well when it comes to economic production in environments where there are not enough strangers to work as a buffer. The same Cambodian or Vietnamese entrepreneurs who are successful in California or the Czech Republic struggle to achieve the same success in their home environment. This is not because they are any less industrious or surrounded by sloth when at home. It is because it is harder to escape the totality of obligations up close.
James C Scott observed that small shopkeepers in many parts of the world are often strangers (often of different ethnic or linguistic origin) who are not tied to local structures of kindship and obligation. It is impossible to run a small shop if it is impossible to refuse to simply give food to people with whom you have a strong bond. This is certainly true in Timor-Leste where many of the small local enterprises struggle.
This is common around the world, so much so that many of the small loan arrangements that have become popular in the development area function less as a way of advancing capital and more as a way of putting capital out of reach of kinship obligations. As another exmple, Leo Howe reports that many of the people working in the hospitality industry in Bali are actually Javanese because the native Balinese cannot be relied on to be always available. This is not because they are ‘unreliable’ in some essentially flawed way but because their obligation to family (often ritual) is too great to suspend through a contract with strangers. Marshall Sahlins’ essay on ‘Stranger Kings’ shows that this applies even to choosing to submit to a ruler.
The strong ties with kin and weak ties with non-kin can cause even greater problems which is why we find such elaborate hospitality rituals around the world. Tourists often misinterpret these under the heading of ‘oh, the people are so friendly here’ but in fact, this is a function of the culture not having a norm around dealing with strangers that does not rely on the notion of guest as temporary kin.
The ancient Greek concept of ‘xenia’ - hospitality to strangers can be very illustrative. We know the root from ‘xenophobia’ but the Latin equivalent gave us both hospitality and hostility. Xenia not only dictates extraordinary measures to take care of guests but it also strictly regulates the behavior of those same guests. Breaking the rules of hospitality has been the source of many problems from the Trojan War to the blood feuds in modern Albania - for instance, as described in Kadare’s “Broken April”.
These are not inevitable consequences of the sort of groupings Brooks is describing and advocating for. Not is he unaware of potential internal problems. But when the entire world is structured through strong kin-like ties, we have not just created an archipellago of utopias - as some in the Seasteading movement seem to imagine - we have a world that is fundamentally different from what we know. The demands of the trade-centred world dependent on industrial production and industrialised aggriculture cannot be entirely ignored in this vision.
## Conclusion: The kin, the strangers, and the state
In conclusion,these rough sketches are not meant to diminish Brooks’ contribution. If I were asked to come up with an alternative to the nuclear family, it is almost exactly, what I would propose. Many ethnographic accounts of impoverished communities all over the world show that these networks often emerge organically. And we should be doing as much as possible to normalize them. They are not some lesser, last resort alternatives to the nuclear family. They are both natural and can be healthy and robust. But they don’t exist in a vacuum and are themselves not static nor do are they uniformly idyllic. There are cracks the size of valleys between them and sometimes even within them. And it’s very easy for individuals to fall through those.
That’s where we still must see a role for the state to protect both individuals and groups from falling to the ground without a safety net as well as adjudicate the parameters of their encounter. At the moment, the state support structure is entirely structured around the schemas and scripts of the nuclear family on the one hand, and the contract between strangers on the other. It needs to recognise a wider range of support networks and obligations which is only possible if we reframe the notion of family. Such reframings are always long and do not progress in a linearly predictable fashion. Brooks’ essay is an undeniably valuable contribution to this process. So despite any quibbles and caveats, I’m all for it.
## How to actually write a sentence: The building blocks of written language
Date: 2020-02-01
URL: https://metaphorhacker.net/2020/02/how-to-actually-write-a-sentence-the-building-blocks-of-written-language/
Categories: Education, Linguistics, Extended writing, Writing
Tags: featured
Some time ago, Thomas Basbøll followed up his excellent post on how to write a paragraph (https://blog.cbs.dk/inframethodology/?p=2676) with a much more daring endeavour on how to write a sentence (https://blog.cbs.dk/inframethodology/?p=2699). And while the post is a pleasure to read, I think it did not quite overcome the challenge the author stated at the start:
“it is substantially more difficult to explain what one does when one writes a sentence than it is to explain what one does when one composes a paragraph.”
Indeed, it is much more difficult to talk about the mechanics of writing the sentence because we generally want to forget we are composing a sentence, whereas we want to focus on the fact we are composing a paragraph. In this, writing a sentence is much like riding a bicycle. You cannot really do it successfully while attending to every aspect of the process. Basbøll’s metaphor here is very apt:
"it’s easier to give you directions to City Hall than to explain how your legs work. Sentences, we might say, are to paragraphs as taking a step is to going somewhere. It’s only once we pay attention to it that we realize how subtle and how stylish such a simple thing can be."
The problem with his solution, though, is that it only focused on the role of the sentence in the process of expressing ideas rather than the mechanics of putting a sentence together. This is because a sentence is an artificial construct. We think of it as a natural unit but, in fact, it is only an accident of history that we’ve started dividing chunks of text with full stops and beginning them with capital letters.
The sentence is just one way of articulating a thought. It could be a list. A phrase. Or a whole stream-of-consiousness story. But through conventions, we think of all of these as inferior kinds of writing. Expressing ourselves ‘in complete sentences’ has been agreed to be the hallmark of educated expression. And whether we agree with it or not, sentence is what we’re stuck with.
## What is a sentence?
There is much debate in linguistics as to what is the foundational building block of language. It could be a phoneme (sound), syllable (much more natural in speech), word (unit of meaning), utterance or text (one chunk of speech with a message). It could also be a phrase. But by far the best candidate is a clause - a unit with one predicate and one subject - even if it is not always easy to define exactly what predicates and subjects are. But whatever the basic building block of language may be, sentence is definitely not it. It’s not even a unit in conversational speech but despite its visual significance, it is not really the basic building block of written language either.
This is because the boundaries of a sentence are completely arbitrary. They are simply there for the convenience of visual processing. The preceding 2 sentences could just as easily have been one. And many people would insist that they would be better as one and then argue over the proper rules of punctuation.
The real problem, and the one Basbøll is actually writing about, is how to express one’s thoughts through writing in a way that generates mental representations in the mind of the reader that are as close as possible to those of the writer. He illustrates it nicely with a quote from Orwell:
"As George Orwell pointed out many years ago, a great deal of bad writing comes out of stringing words and phrases together that are completely unrelated to any pictures that might form in any human being’s head."
There is something in this. We might argue that at least what is written represents what is in the writer’s head. But often our written words are just an echo of what was in one’s mind rather than a rendering of a mental image. Who has not had the experience of reading something they have written and not being completely certain what they meant by it?
So, making sure you build the right image in the reader’s mind with your words is excellect advice. But where Orwell, Basbøll’s essay and many others come up short is in explaining how to go about stringing those words together in just the right way so that they can trigger the right image in the reader’s mind. In this post, I’d like to suggest some ways in which we actually may go about learning to write a sentence to achieve this aim.
## Dual articulation, riding the bike and Krashen’s monitor
But before we go any further, let’s look a bit more closely at the nature of the difficulty identified by Basbøll. That is: What we really want is to express ideas, not craft sentences. We want to go effortlessly from idea to sentence or better still from idea to paragraph. But we have to pass through many intermediate steps before we get there. Choosing words, calling up their spelling, deciding on their relative placement, whether we should add any endings, and then telling our fingers to type them. It’s even more complex in speech, where we have to arrange our mouths, tongues and teeth into complex configurations and coordinate all of that with the work of the lungs and the epiglottis.
In other words, before we can articulate a thought, we have to articulate a lot of other things. This has been called the ‘dual articulation’ of language. Dual articulation is one of the most underappreciated aspects of language. It is what makes non-native language learning so hard. And writing is certainly not native to any of us.
We spend a lot of time trying to learn all the rules of articulating words and sentences. But in order to successfully and fluently articulate ideas (which is what language is there for after all), we have to make the complex process of articulation of all the building blocks of language disappear. If we were to attend to all aspects of it, we would be permanently tongue-tied.
This is an experience that any learner of a foreign language has had when trying to use their newly acquired knowledge outside the classroom. Stephen Krashen has proposed the monitor hypothesis where the goal of language acquisition is to reduce the role of the grammatical monitor. In the same way that native speakers not only do not pay attention to how they put words and sentences together, learners must get rid of this additional burden. Speaking a language then is just like riding a bike. If you pay attention to all the tiny movements that are involved in peddaling while keeping balance, you fall off. But equally, if you miss any of them out, you fall off, as well.
So what are we, who want to teach others to write sentences, to do? On the one hand, we have to tell them about the principles of sentence structure that they were not able to suss out from their own reading. But on the other hand, we have to lead them to completely forget about all of them when they most matter and just write.
## Writing as editing and editing as reading
Luckily, writing is not as ephemeral and fast flowing as speaking. We can always come back to a sentence we wrote and change it beyond all recognition. So, to teach somebody how to write is really teaching them how to edit. And a big part of teaching somebody how to edit, is to teach them about what to pay attention to when reading.
To be clear, a fluent writer can formulate a sentence without much need for further editing. But editing is a process through which such facility can be acquired. And even the most expert writers will need to come back and edit some of their sentences.
What does an editor pay attention to? They will tell you that they look at two things: 1. does the sentence make sense and 2. does it flow from the previous sentences and into those that follow. They will also look at more formal aspects such as spelling, undue repetition of words, stylistic appropriateness, etc. But 1 and 2 (sometimes also called coherence and cohesion) are the fundamental structural jobs a sentence has to perform.
## How to craft a sentence
This finally brings us to the ultimate aim of this post. How to actually put a sentence together. This is, of course, impossible to cover in a single blog post. There are shelves in libraries around the world groaning under the weight of volumes that barely scratch the surface of all the aspects of a well-crafted sentence. Yet, people have managed to become competent or even admired writers despite all that. So, there must be way.
### Learning to craft a sentence
It is important that aspiring writers think about the learning process as much as about the actual components of a sentence. And the process is very simple:
- When you read something, spend at least some of the time, looking at how it is put together. If this is not what you naturally do, set aside some time to do this as part of your reading.
- Form hypotheses about the rules the author used and then try them out yourself. It does not matter whether these hypotheses are correct ‘grammatical’ rules or even whether they look like grammatical rules. It just matters that you can do something with them.
- Leave what you wrote sit for a while and then come back to it. Read it again and see if it still makes sense. Then go back and look at how what you wrote differs from what you intended. And also compare this with other writing.
- Read things out loud or have them read to you (e.g. by text to speech). This will sometimes allow you to notice things about the text that you may skip over when reading silently.
- Do this a lot.
With that in mind, let’s finally have a look at some of the things you have to know about how to write a sentence.
### Making a sentence make sense - Coherence
For a sentence to fulfil its ideational function, it has to make sense. This means that the sentence must not only contain the idea you want to express, it must not get in the way of that idea. When you’re editing your sentences to make sure they make sense, ask yourself these questions:
- Is it possible to read the sentence in other ways? Sometimes, when a sentence comes out of your head, you are blinded to its other possible meanings. Read it out loud, or ask your software to read it out loud for you.
- Have you chosen the right words? This seems obvious but choosing the words that mean what you want to say is not a given.
- Are the subjects of the clauses linked clearly to their verbs? Or, is it clear what the verbs in your sentence are describing? Conversely, is it clear what is happening to the nouns in your sentence? A simple test is to try to reduce the clauses in your sentence just to underlying verb and noun pair (or subject and predicate). Then keep adding the other words until the sentence is back together. If this sounds like old-fashioned parsing, it’s because it is. But sometimes it is necessary to strip your sentence bare and then slowly add only the necessary components back. Often it is the only way to make a sentence that got away from you make sense again.
- Have you compressed too much into a single sentence? Can you expect that your readers have the same background and can take a hint?
- Is it clear what the pronouns refer to? When you’re writing, your subject is very active in your head. So, it is very common to keep using pronouns or other vague words to refer to what you’re talking about. It is safer to use pronouns a bit more sparingly and repeat more often. While there’s a lot of research in this area, there is no one rule for how to do this right. But most of us were warned against repetition by our teachers, so a good rule of thumb is to repeat a bit more often than you feel comfortable.
- Have you used the keywords in the right context? Sometimes words have multiple meanings and the one you are trying to express may not be the one most readers associate with it. Perhaps the best tool to help you here is a corpus. The iWeb corpus (https://www.english-corpora.org/iweb) is a great tool for checking how words are used.
### Making a sentence hold together - Cohesion
But even if your sentences make sense and use all the appropriate conventions, they still have to hold together and fit in with the rest of the text. This is often the easiest problem to overlook because you have an overall picture of the text in your mind, so it all flows perfectly in your head.
But your reader will have to build a picture of the text from scratch. And, also, they may not always read perfectly linearly, so even a sentence read out of context should make it clear where it relates to what came before.
Here are some questions to ask yourself when you’re editing a sentence:
- Have I made the right logical connections? If one thing is caused by another, is there a ‘because’ or a similar conjunction to make the link explicit?
- Have I not put too much distance between closely related things? Long parentheticals can be fun but make it very easy for the reader (as well as the writer) to get lost.
- Have I focused the reader on the right point? The topic (or known information) of a sentence is usually at the beginning and the focus (or new information) should come at the end.
- Have I given the reader too much work to parse the sentence? If so, can I make it easier by splitting the sentence into shorter chunks?
- Can I move some things to a later sentence?
- Have I expressed a clear link to what came previously?
- Have I placed the sentences in the right order? Don’t be afraid to move a sentence to the end to make sure the key information comes earlier.
A useful tool to use here is the Hemingway Editor. It will highlight sentences that are too long. Now, in many genres, such as academic writing, long sentences are not always a problem. They’re almost the expectation. But a sentence that goes on too long should be a signal to you, that you may not have expressed your idea clearly. I find that my long sentences are often just piles of ideas that need to be taken apart and given more air.
### Making a sentence communicate what you want how you want it: Genre and style
Even if your sentence makes sense, your reader must be willing to try to read it. This means that you must meet as many of their expectations as possible so that they can focus on the meaning. You do this by conforming as closely as possible to the conventions of the genre you work within. If you do break these conventions, make sure you’re doing it for a reason.
If you’re writing an academic essay, stay within the [register] of academic language. This is where the various guides on academic English come in. They break down language into communicative functions like argumentation, persuasion or disagreement. And then they give you lots of appropriate phrases to achieve that function.
This is also where you need to do a lot of targetted reading in the area you want to write in. Don’t just read for content, read with an eye on the way people express themselves. Narrow your area as much as you can.
For example, there’s not just one ‘academic English’. Each little subdiscipline has its own conventions, so it’s worth paying attention to those. One piece of advice given is, before you submit a paper to a journal, read other papers that had already been published there. They will give you a clue as to the expectations. This applies at all levels, not just the sentence.
There are technical tools that can help you. For instance, you can paste your text to the Analyze tool on AcademicVocabulary.info (https://www.wordandphrase.info/academic/analyzeText.asp) and check the words you used against a corpus of academic writing.
### Writing and editing process tips
Finally, here are some tips about the process of writing and editing your text at the level of a sentence.
- Don’t edit every sentence independently - only edit when you’ve written several of them to make sure they hang together.
- Feel free to delete a sentence. Often, once we’ve written something, we feel possessive about it. But often, deleting something can be very helpful. Like pruning a tree.
- Feel free to split a sentence in two or three. Sometimes, it will give you space to express yourself more clearly. But sometimes, it will just give your reader a visual cue that a new idea is coming. Or at least some space to take a breath.
- By the same token, don’t be afraid to start or end a sentence with a preposition or a conjunction. It’s much better than twisting yourself around.
- Don’t be too scared of long sentences. Sometimes, joining two shorter sentences together makes the text flow better.
## Reflections and conclusions
The abiding concern of anyone telling somebody else how to write is whether they themselves measure up to what they preach. Or at least, it should be. We know that Orwell used more passives than average (http://www.lel.ed.ac.uk/~gpullum/passive_loathing.pdf) while advising against them, Strunk and White (http://ling.ed.ac.uk/~gpullum/50years.pdf) used many of the same constructions they advised against, and the Plain English campaign proponents (/2012/09/the-complexities-of-simple-what-simple-language-proponents-should-know-about-linguistics/) don’t always use simple language.
Equally, I cannot guarantee that every sentence in this guide is a paragon of what a well-crafted sentence should be. I know my limits. I tend to write more than needed and not cut out enough having learned my English syntax at the feet of PG Wodehouse. But demonstrating perfection at the level of the sentence is not the point of this post, and neither should that be the aim of most writers. The aim is to get the point across and then to move on.
But the most important conclusion is that hesitant writers must pay attention to the learning process. It is not possible to explicitly follow all the tiny little rules for putting together a sentence. You must internalise the shapes and bigger chunks, so that you can focus on experessing your ideas. This can only be achieved through deliberate practice. And editing what you wrote is the most crucial part of that practice. Great writers have great editors, or if they're poor, they're their own great editors.
Image by Free-Photos (https://pixabay.com/photos/?utm_source=link-attribution&utm_medium=referral&utm_campaign=image&utm_content=1209121) from Pixabay (https://pixabay.com/?utm_source=link-attribution&utm_medium=referral&utm_campaign=image&utm_content=1209121)
## Potemkin wisdoms, phronesis and Pixar: How wise sayings protect us from meaning
Date: 2020-01-21
URL: https://metaphorhacker.net/2020/01/potemkin-wisdoms-phronesis-and-pixar-how-wise-sayings-protect-us-from-meaning/
Categories: Knowledge, Metaphor, Philosophy, Extended writing
Tags: featured
## TL;DR
This is an exploration of the difference between wisdom and practical wisdom (phronesis) triggered by this quote from a talk by Ed Catmull (https://www.google.com/url?q=https://www.youtube.com/watch?v%3Dk2h2lvhzMDc&sa=D&ust=1579646399857000):
“Once one can articulate an important idea into a concise statement, then one can use this statement, and not have to have the fear of changing behavior.”
The main lesson is: if we confuse understanding with repeating its summary, we hollow out its meaning and can no longer rely on it to inform what we do.
It explains why adopting even great advice often does not result in success. It explains why most charismatic reforms fail when spread out more widely. It explains why adopting even Catmull’s advice may not make you into Pixar.
## Why are advice books so often free of content?
Ed Catmull (https://www.google.com/url?q=https://en.wikipedia.org/wiki/Edwin_Catmull&sa=D&ust=1579646399859000) was an engineer suddenly put in charge of a company, so he did what engineers do. He went looking for a manual on how to run a business. In his words (https://www.google.com/url?q=https://www.youtube.com/watch?v%3Dk2h2lvhzMDc&sa=D&ust=1579646399859000) (slightly edited from transcript with punctuation inserted):
“I had to learn a lot about business quickly and I hadn't gone to any school so I just read a lot of books and there were a few bits and pieces but I got to say, for the most part, I didn't get a lot of traction with them.”
His solution:
“So I said, well maybe the problem is I just got to get to the essence of them. There's a service that gets the summaries of business books so I tried that. And that was actually an amazing experience, because, in reading the essence of the books, I realized they were content-free.”
But maybe the problem was not in the business books themselves:
“What's going on here? Is it the fact that the book doesn't have any content, which is probably true in many cases, or is it that you can't take some of these things and reduce them in a meaningful way?”
## When it’s more important to say than to do
But Catmull was not interested in the nature of understanding. What bothered him was that the principles, once reduced to a slogan, made it impossible to determine practical success. Or rather, that you couldn’t tell who was actually good at something by the principles they espoused. It started in the film industry where the accepted wisdom was that ‘story is the most important thing’ about a film, but as Catmull discovered:
“... every studio says the same thing. Everybody says the stories are the most important thing, even if the story was drivel. It might be true, in fact it is true, but it doesn't affect behavior. … It’s one of those things that is true and you agree it is true and you say it, but that doesn't mean anything.”
He found the same thing in architecture when he worked on building projects. Everybody agrees that you should design buildings “inside out” but that applies to architects of great buildings as well as the ones of awful buildings. On its own, this may be an observation many people have made - the Dilbert cartoon series is based precisely on the fact that we all recognize what people talking in empty phrases sound like. But Catmull’s formulation of it is very striking and less glib:
“The phrase [designing buildings inside out] is important to this community, it just does not have any effect on behavior.”
Early on in his talk Catmull asked what is more important “good ideas or good people”, this seems to point in the direction of ‘good ideas’ not being very important if they don’t help people do better things. But it also shows that ideas are tied to their expressions and those expressions play many more roles in their communities than communicating what the ideas are about. They tie their communities together through the ritualistic profession of creeds that signify belonging. It is more important that people ‘agree with’ or ‘proclaim’ the phrases than embody them through actions. (This, of course, has a long history in religious reform. Or educational reform. I recently gave a talk exploring how the term ‘pedagogy first’ (https://www.google.com/url?q=https://altc.alt.ac.uk/online2019/sessions/359/&sa=D&ust=1579646399861000) is mostly absent of meaning on which actions can be based.)
## Compressing ideas renders them meaningless
None of this will be news to anybody who has ever worked for a big institution (University, corporation, Government department). Catmull expressed this very starkly in what I would consider a key quote from his whole talk:
“Once one can articulate an important idea into a concise statement, then one can use this statement, and not have to have the fear of changing behavior.”
This can even be used very strategically. In my research on personalisation (https://www.google.com/url?q=https://www.researchgate.net/publication/277329463_Thinking_talking_and_writing_computer_programmes_about_personalisation_at_City_College_Norwich&sa=D&ust=1579646399862000) I ‘discovered’ that people were quite strategically looking through all the things they were already doing and trying to label them as ‘personalisation’. Often, in conversations about how to apply new methods, more time is spent on ‘labeling’ different activities than thinking about what to do. Catmull’s experience is the same:
“I see this over and over again. I can summarize some of these things, but the real issue is what do we do?”
But this is not just an issue with ideas being diluted through institutionalisation. It is a cry for help about the very possibility of making ideas mean anything at all. What good are great ideas if nobody can do anything with them. Many of the ideas “we all agree with” are expressions of genuine ‘wisdom’ but by the process of spreading them, we hollow out their content. And what we end up is wisdom painted on top of an empty box. We even have a story about this: ‘The Emperor’s New Clothes (https://www.google.com/url?q=https://en.wikipedia.org/wiki/The_Emperor%2527s_New_Clothes&sa=D&ust=1579646399862000)’.
## Potemkin village, Potemkin wisdom
But an even better story than the one about a naked Emperor is the one about General Potemkin and his villages.
I grew up in a communist dictatorship and calling something a Potemkin’s village (https://www.google.com/url?q=https://en.wikipedia.org/wiki/Potemkin_village&sa=D&ust=1579646399863000) was common when reflecting on the propaganda of the regime. The phrase referred to a Russian general who created fake facades (https://www.google.com/url?q=https://en.wikipedia.org/wiki/Facade&sa=D&ust=1579646399863000) on buildings in villages so that the visiting Empress (https://www.google.com/url?q=https://en.wikipedia.org/wiki/Catherine_the_Great&sa=D&ust=1579646399864000) would think that they were prosperous and all was well in her realm. These facades could be moved from village to village, so as the Empress and her retinue travelled around, they could appreciate the quality of her government.
Pithy summaries of great ideas are like facades of beautiful buildings. They first appear as expressions of joy at the greatness of the structure inside. But unless we actually go in, walk through the rooms and corridors, or even better, live in them for a while, the greatness of the building is just an assumption. But we don’t always have the time to go in all the great buildings we see, and we definitely don’t have the time to live in them. So, when our own building is crumbling inside and out, we find it much easier to just paint the outside of it. Then anybody walking by will think it’s a great building and maybe we’ll even convince ourselves that the building is great.
Thus we create Potemkin’s wisdoms. Slogans we take from from one situation to another and paint them over whatever is actually happening. Unlike Potemkin, we don’t do it to deceive the Empress. From the outside, we see the beauty of the building, the truth of the idea. We want to embody that truth, so we paint it on top of what we do and admire it from afar. We are deceiving ourselves. But after a while the reality shines through the peeling paint and we go out looking for a new facade. This is the cycle of the hidden utopia.
## The fundamental paradox of understanding
The reason this resonated with me to the point of tracking down the transcript of the talk is that this is a topic that I’ve been grappling with for almost 30 years. I’ve been returning to the question of understanding complex issues through summaries ever since I heard the Czech philosopher Peter Rezek (https://www.google.com/url?q=https://cs.wikipedia.org/wiki/Petr_Rezek_(filosof)&sa=D&ust=1579646399865000) ask why did philosophers write these long books when we can then just talk about them in what is essentially aphorisms? Rezek’s answer (if I remember it correctly) was that we need to read the complete books, but he also wanted to explore the underlying tension.
The tension is that we (and this is me, not Rezek, speaking) can only really access the content of great books retroactively through reductions our mind creates in the process of understanding. Of course, the process of reading the book also changes the conceptual landscape of our mind against which any understanding is viewed. But when we try to recall that understanding and employ it in further thinking, we draw on those aphorisms and hope that the accompanying change in landscape contains all the important components to fill the pithy phrase back up with meaning. But often that is not that case. The time of use what was understood and the time of applying that understanding are far removed. But even if they were not, the understanding was always partial to begin with.
We also need to distinguish between understanding as the ability to draw the same inferences as the author of a text versus understanding as a moment of enlightenment. Enlightenment is a single event, but understanding is a process (https://www.google.com/url?q=/2019/05/explanation-is-an-event-understanding-is-a-process-how-not-to-explain-anything-with-metaphor/&sa=D&ust=1579646399865000). We are often almost ecstatically aware of the moment at which we finally ‘got’ something. But the tedious process of developing the kind of understanding (https://www.google.com/url?q=/2019/06/5-kinds-of-understanding-and-metaphors-missing-pieces-in-pedagogical-taxonomies/&sa=D&ust=1579646399866000) we might actually do something useful with happens largely under the radar of our consciousness. We see understanding and explanation in charismatic terms but the actual achievement of it is a matter of routine (https://www.google.com/url?q=https://en.wikipedia.org/wiki/Charismatic_authority&sa=D&ust=1579646399866000).
## Wisdom contra Phronesis
And then, Catmull adds the dimension of ‘understanding as a social obligation’. This is very much reminiscent of the pragmatist (https://www.google.com/url?q=https://en.wikipedia.org/wiki/Pragmatism&sa=D&ust=1579646399866000) notion of truth. We only signal understanding through action in front of our peers. And by far the easiest action is that of repeating an accepted Shibboleth. Elezier Yudkowski has aptly called the sort of thing that happens in school guessing the teacher’s password (https://www.google.com/url?q=https://www.lesswrong.com/posts/NMoLJuDJEms7Ku9XS/guessing-the-teacher-s-password&sa=D&ust=1579646399867000). When Catmull says “this phrase is important to the community,” he is talking about such a password. But, then, he observes, “it has no impact on behavior”. What he is after is what the ancient Greeks called “phronesis (https://www.google.com/url?q=https://en.wikipedia.org/wiki/Phronesis&sa=D&ust=1579646399867000)”, a practical wisdom. The sort of wisdom one can only gain through experience informed by knowledge. But this kind of knowledge cannot be expressed through a summary. Or perhaps not even communicated at all.
The Greeks helpfully differentiate between sophia (https://www.google.com/url?q=https://en.wikipedia.org/wiki/Sophia&sa=D&ust=1579646399867000) (wisdom), phronesis (https://www.google.com/url?q=https://en.wikipedia.org/wiki/Phronesis&sa=D&ust=1579646399868000) (practical wisdom), episteme (https://www.google.com/url?q=https://en.wikipedia.org/wiki/Episteme&sa=D&ust=1579646399868000) (knowledge) and techne (https://www.google.com/url?q=https://en.wikipedia.org/wiki/Techne&sa=D&ust=1579646399868000) (skill, craft) - although this was a lot more complicated (https://www.google.com/url?q=https://plato.stanford.edu/entries/episteme-techne/&sa=D&ust=1579646399868000)with (https://www.google.com/url?q=https://en.wiktionary.org/wiki/%25CE%25B5%25E1%25BC%25B4%25CE%25B4%25CE%25BF%25CE%25BC%25CE%25B1%25CE%25B9&sa=D&ust=1579646399869000) many (https://www.google.com/url?q=https://en.wikipedia.org/wiki/Gnosis&sa=D&ust=1579646399869000) more (https://www.google.com/url?q=https://en.wikipedia.org/wiki/Nous&sa=D&ust=1579646399869000) distinctions (https://www.google.com/url?q=https://en.wikipedia.org/wiki/Poiesis&sa=D&ust=1579646399869000) floating about. So if we are after the sort of judgment that comes from phronesis and has to be acquired rather than taught, what of the other forms of knowledge? Do they play no role at all? We certainly want people to know things (episteme) and be able to do things (techne), but do we really need them to be wise (sophia)? What if wisdom is only the ability to make pronouncements that are essentially empty of content on which one could base behavior?
In that case, wisdom is a ritualistic social function. We need wise people to tell us things like ‘measure once, cut twice’ or ‘paralysis through analysis’ to show us what we are as a community. And because these pearls of wisdom fit any situation in one way or another, it is not difficult to find them profoundly true without thinking about their emptiness. They give us the same sense of instantly gratifying insight that horoscopes and star signs do. When we read a description of our star sign, it can give us an almost euphoric sense of recognition of ourselves as being part of a bigger universe. In the same way, reading a wise saying, or a business book, can give us a glimpse of that sense of oneness with the world that gnostics or yogis must feel. We get to feel wise, in the know, with it - we finally got it.
And it is this feeling that Catmull warns against. In the preface to his book, he says:
“What makes Pixar special is that we acknowledge we will always have problems, many of them hidden from our view; that we work hard to uncover these problems, even if doing so means making ourselves uncomfortable; and that, when we come across a problem, we marshal all of our energies to solve it.”
This might be the very definition of phronesis - acknowledgement of the fact that understanding is a process without an end. And that it has an emotional dimension. Understanding is not a question of insight but a matter of practical never-ending work. We can use wisdom - such as ‘understanding is a process’ as a catalyst for the actual work that’s necessary. But on its own, it’s just an ornament that we put on something to make it look better than it is.
## Practical wisdom in the face of the complexity of understanding
But is it enough to read Catmull’s book or listen to his talk? We can follow his advice step by step and still fail in our aims. What happens to his wisdom (phronesis) when applied to his own pronouncements? Do they have any content at all when divorced from their original context? I’ve spent over 2000 words so far exploring their implications and I’m barely scratching the surface. We can say ‘what he says makes sense’ but it is actually us making it sense out of his words and our situation.
The problem is that our intuitions about meaning are wrong (https://www.google.com/url?q=/2018/05/therapy-for-frege-a-brief-outline-of-the-theory-of-everything/&sa=D&ust=1579646399871000). When we reflect on language (metacognition), our reflections take the form of that builds on a schema that very closely resembles a dictionary. In a physical sense, a dictionary puts meanings next to the words that express it. As if words were just little pointers to meaning. At its most schematic, this schema takes on the shape of one to one correspondence: fromage = cheese, dog = [picture of dog]. In certain contexts, we can acknowledge homophony or polysemy: dog = 1. animal, 2. food. At its most sophisticated, we imagine the right-hand side of a dictionary as an encyclopedia.
But that is not how words are used. We do not just put them next to each other and easily combine their meanings. For instance, when we apply adjectives to nouns, we don’t just add up the two meanings as we do with 1 + 1 = 2. Yellow cheese and yellow dog will seem superficially the same - ‘yellow’ + thing = yellow thing. But that is only at the highest level of abstraction. If we actually want language to communicate something somebody can draw some useful inferences from, a lot more has to happen. The yellow of cheese is very different from the yellow of the Sun or a crayon. Also, it is an expected color. But the yellow of a dog is an unusual colour. We can call it blond or something but we do not expect it to be the same as that of cheese. If we see the words ‘yellow dog (https://www.google.com/url?q=https://www.google.com/search?q%3Dyellow%2Bdog%26sxsrf%3DACYBGNTX5He6V148dtbAoPMEqZ2KiOhgqQ:1579502163796%26source%3Dlnms%26tbm%3Disch%26sa%3DX%26ved%3D2ahUKEwiriZSLyJHnAhV2QxUIHVCKDCsQ_AUoAXoECA4QAw%26biw%3D1463%26bih%3D713&sa=D&ust=1579646399872000)’ we may expect a picture of an animal or an animal that has been covered with yellow paint. Also, when we cut a block of cheese in half, we expect the color to be more or less uniform throughout. The vet performing a surgery on a ‘yellow’ dog (whether blond or painted) would be extremely surprised to find the dog to be yellow on the inside.
What does this have to do with anything? Only that, if applying the simple label of ‘yellow’ to simple nouns like ‘cheese’ or ‘dog’ is this complicated, how can we expect to simply apply complicated labels like ‘we acknowledge we will always have problems’ to complicated institutions like ‘Pixar’ or ‘Coca Cola’ or ‘Bob’s Bodega’?
Because of the complexity of meaning making, people understand the same words applied to the same situations differently. This is such a fundamental fact of everyday life as to be considered trivial. So obvious that it barely deserves a mention. But if it is so obvious, why is it completely absent from our basic schema of meaning and understanding? Why don’t we add this to our consideration when we say things like ‘people need to be more considerate’? Our common reaction on hearing something like that is to agree. But instead we should be asking “what do you mean?” What does a ‘considerate person’ look like in your head? What are your schemas and scenarios? In the same way we might ask somebody “what do you mean by ‘yellow dog’?”
But, of course, most of the time we cannot. Conversation would be impossible. Indeed, life would be impossible if we couldn’t rely rough schematic understandings. Instead, we make sense of things as we go along. Sometimes we do ask for clarification, but usually we look for cues, or just nod along and hope it will all make sense in the end. With very complex meanings - such as famous principles - we often just assume an underlying structure that was never there - like the Empress looking at a facade of an empty building.
When I say ‘understanding is a process’ and ‘knowledge is social’, I mean exactly that. We build up meanings as we go along, make assumptions, ask questions, hope somebody else knows what that means, and so on. That is exactly what happens when somebody reads a book like Camull’s - they start building images that they then share with others and hope to come to some sort of understanding. And part of that understanding will be ‘Catmull ran Pixar, his words will make us more like Pixar’ and ‘Catmull only ever ran Pixar, what does he know about my neck of the woods’ and ‘I just want this meeting to be over so I can go get lunch’.
Catmull himself was aware of the problem and instituted processes at Pixar that tried to break up the routines of meaning making and trigger less schematic reflection. But what happens if somebody who is not Catmull - doesn’t have his style, his conviction, his understanding of what he actually means - in short his ‘charisma’? We are back to Dilbert! All through history, charismatic reforms have failed when applied across larger numbers of peoples and institutions. The followers of St Francis of Assissi soon start acquiring possessions in his name. Revolutionary leaders fighting to overthrow oppression soon become the oppressors. Teachers fired up by philosophies of empowering the child soon start spending most of their time taking attendance and marking homework.
Unfortunately, it’s even more complicated than just people misunderstanding the wisdom of those imbued with the virtue of phronesis. We cannot even be sure that those who have the requisite practical wisdom even understand it themselves. They may be telling us about it but may be misdescribing their own understanding by applying idiosyncratic interpretations to common schemas. Or sometimes they’re just saying stuff to justify something that happened mostly outside their control.
I once heard an executive bragging that he turned around a failing company by making everybody count the number of paperclips they were using. This, according to him, got everybody to focus on costs and led to a turn around in the company’s balance sheet. On the surface, this is a plausible story, and we even have common sense wisdom in the form of ‘look after the pennies and the dollars will look after themselves’. But on second’s reflection, it is utterly implausible. He and lots of others most likely did a lot of other things and the stupid counting of paper clips just got in the way. We can imagine another executive coming in and saying ‘focus on what’s important’ or as the folk wisdom has it ‘don’t sweat the small stuff’. Just like Steve Jobs did when he came back to Apple.
So we should be skeptical of Catmull’s sine qua nons in his talk and in his book. Lots of successful and creative companies succeeded doing exactly the opposite of what he recommends. But even if his prescriptions may only work sometimes, his description of what happens, when we brandish slogans as shields, is still the best we got. His solution is to be a sort of every day philosopher, and he’s not the only one to advocate for this - Richard Rorty (https://www.google.com/url?q=https://plato.stanford.edu/entries/rorty/&sa=D&ust=1579646399875000) and John Elliott (https://www.google.com/url?q=https://people.uea.ac.uk/john_elliott&sa=D&ust=1579646399875000) are two others I can think of. But practical philosophy is too easily confused with professional philosophy - thinking pretty thoughts and expressing them in pretty words.
However, just declaring ‘everyone should be a practical philosopher’ is not enough. Philosophers may ask questions like ‘what do you mean by that’, but because most of the time, they have those same intuitions about meaning as something you find in a dictionary, these questions lead nowhere. The philosopher’s love of wisdom is often the love of a lexicographer enamoured of putting words next to their definitions. These sorts of questions only have value if what we seek is phronesis and not sophia. Practical wisdom, judgement in the context of actual collective endeavour. Unsatisfactorily, it is a process without an outcome. It can never end in a dictionary entry which can at best capture it frozen in time. And, if that is that we end up with, then all we have is a phrase we can carry around with us as a shield to protect us from actually changing what we do. Or to close with Catmull’s words:
“Once one can articulate an important idea into a concise statement, then one can use this statement, and not have to have the fear of changing behavior.”
## Background
Many years ago I heard John Siracusa (https://www.google.com/url?q=http://hypercritical.co/&sa=D&ust=1579646399876000) on a podcast (https://www.google.com/url?q=https://5by5.tv/hypercritical&sa=D&ust=1579646399877000) say something like “nothing hides problem like success” when talking about issues with Apple’s software when the company was hitting $1 billion in market value. It immediately made sense to me. It is hard to try to give advice to people who are successful at something. But recently I had the urge to track down the origin of the quote and after a bit of searching I discovered (https://www.google.com/url?q=http://5by5.tv/hypercritical/12&sa=D&ust=1579646399877000) that John got it from Pixar’s then President Ed Catmull’s 2009 talk at Stanford (https://www.google.com/url?q=https://www.youtube.com/watch?v%3Dk2h2lvhzMDc&sa=D&ust=1579646399878000). I watched it and realised that it is relevant much more broadly than just as a source of one pithy phrase. In fact, it had at least 2 important insights.
- Success hides problems
- When powerful ideas are encapsulated in pithy sayings, they lose their power to change behavior
I went looking for the first but it’s the second one that was the inspiration to write the above.
## Postscript
I found it interesting that the summary of Catmull’s talk (https://www.google.com/url?q=https://www.gsb.stanford.edu/insights/ed-catmull-we-constantly-seek-out-small-crises&sa=D&ust=1579646399879000) published by Stanford where he gave this talk skirted the first and completely ignored the second insight. But without the second, the first one is meaningless. It will just sit there being numbly repeated by all, nodding along with its wisdom and doing nothing about it.
I had a quick look at Catmull’s book Creativity, Inc (https://www.google.com/url?q=https://www.amazon.com/Creativity-Inc-Overcoming-Unseen-Inspiration/dp/0812993012&sa=D&ust=1579646399880000), and he elaborates on the second point in chapter 3 with more examples and metaphors. The metaphor I found most apt is the one of a suitcase but I only read it after I finished this post so it was too late to work it.
“Imagine an old, heavy suitcase whose well-worn handles are hanging by a few threads. The handle is “Trust the Process” or “Story Is King”—a pithy statement that seems, on the face of it, to stand for so much more. The suitcase represents all that has gone into the formation of the phrase: the experience, the deep wisdom, the truths that emerge from struggle. Too often, we grab the handle and—without realizing it—walk off without the suitcase. What’s more, we don’t even think about what we’ve left behind. After all, the handle is so much easier to carry around than the suitcase.”
## So you think you have a historical analogy? Revisionist history and anthropology reading list
Date: 2019-09-08
URL: https://metaphorhacker.net/2019/09/so-you-think-you-have-a-historical-analogy-revisionist-history-and-anthropology-reading-list/
Categories: History, Scholarship, Extended writing
Tags: featured
## What is this about
### How badly we’re getting history
While the world of history and anthropology of the last 30-40 years has completely redrawn the picture of our past, the common perception of the overall shape of history and the development of humanity is still firmly rooted in the view that took hold in the 1800s’ mixture of enlightenment and romanticism.
On this view, we are the pinnacle of development, the logical and inevitable outcome of all that came before us. The development of what is us, the changes in history and culture, can be traced in a straight line from the primitive of the past to the sophisticated of the present. From the savage to the civilized (even if we may eschew these for more polite terms).
But nothing could be farther away from the truth. The shape of global history looks nothing like what we have in our minds from textbooks and popular culture. For a start, it is a lot more complicated, circuitous and fuzzy than we might imagine. That won’t surprise many people. Things are always more complicated when looked at closely. But what I would suggest is that the popular image has completely misplaced the centre of gravity of historical and cultural development. It is the universe before Copernicus and Gallileo, it is the physics before Einstein and Heisenberg.
Yet, all we need to find the right balance is readily available in print, online lectures and courses. We just need to seek it out.
### What is on this list
In this post, I compiled what I consider key books of the last 20 years (with a few older exceptions) that can help anyone get a better picture of the history of human politics and culture. And through that history, we can also see the balance of the present better.
Not all these books are flawless and they all bring new biases into the picture. No doubt, they too, will eventually be subject to revision as new perspectives open up. Also, they don’t entirely reject all that came before them. They simply provide a better balance and shine light in important blind spots.
I can imagine that many people reading any one of these books might feel compelled to reject them as outliers. But together, they are hard to ignore. They come from different perspectives and disciplines, yet, they complement and reinforce each other.
This was originally meant to be a short list of a few key works but as I was going through my notes, I kept adding new ones. I tried to keep the list to books that synthesize larger areas rather than histories or ethnographies of individual societies even though, these can often be as illustrative.
Most of these books are histories or contain historical data. Yet, many are written by anthropologists or historians with a distinctly anthropological point of view. This very much reflects my personal bias towards the ethnographic.
I divided the list into 2 sections: 1. Easy reads for a general audience and 2. Dense and extensive works for specialists. But in this, I was very much going by intuition.
I decided to provide some illustrative quotes for each book but I went a bit too far with some of them. At the same time, I could have quoted many more important passages. Remember, they all make much more sense in context.
Where available, I also provided links to podcasts or online lectures by the authors. I also compiled a YouTube playlist with key videos (https://www.youtube.com/playlist?list=PLEl9d2qvKkBs95yHG62Q5LIjj9mWjaI2u) which I will keep up-to-date as I discover more.
I would also recommend to anybody that they listen to the New Books Network (https://newbooksnetwork.com) podcasts. I find those from the New Books in History, Milirary History, South Asian Studies, Islamic Studies, Anthropology and Genocide studies particularly illuminating and would recommend that anybody goes through the archive, as well.
### What are the key lessons
This section was rewritten based on Reddit comments. (https://www.reddit.com/r/slatestarcodex/comments/d1gbsj/so_you_think_you_have_a_historical_analogy/)
The overarching message of these books is one of anti-reductionism. They do not look for inevitable overarching trends but they do show repeating patterns. The key points that stand out to me as a lesson to take away from reading these books could be:
- The global dominance of Western-European culture and politics is a lot more recent than our history books taught us pretty much starting with the Industrial Revolution and not completed until the end of the 19th century.
- The balance of global history lies in the East rather than the West. Even those we consider the roots of our civilisation (Rome, Greeks) looked to the East.
- We are blinkered by focusing our perspective on civilizational artifacts such as architecture and writing. This leads us to overlook important political and social units that outnumbered those we can see at any one point in history.
- The role of the state throughout history was much more complex and uncertain than it may seem from today’s perspective. It was much weaker, less stable and more transient. And it was also not nearly as attractive to its subjects - ie. walls were often built more to keep people in than out.
- We cannot view the ‘hunter gatherers’ and other ‘pre-technological’ societies of today as remnants of previous evolutionary stages of history. They are as much part of modernity as the technologically-dependent urban centres we know.
### Who should read this
- Anyone who thinks ‘Guns, Germs, and Steel’ is the last word in historical analysis.
- Rationalists, economists and futurists. I very much enjoy listening to podcasts like EconTalk (http://www.econtalk.org) and Rationally Speaking (http://rationallyspeakingpodcast.org). But whenever they or their guests make any points regarding history, I cannot but cringe.
- Anyone who makes historical analogies based on what they learned in school.
- When I last worked with Peace Corps volunteers, I shared some of these books with them and they were well received. So I think many development and international policy workers would also benefit.
- Curriculum reformers in the mould of Michael Gove (https://www.politics.co.uk/comment-analysis/2013/05/09/michael-gove-s-anti-mr-men-speech-in-full) or Pat Buchanan (https://en.wikipedia.org/wiki/The_Death_of_the_West).
## Easy, accessible reads
I felt the books in this section are more accessible and aimed at audiences outside the strict confines of their discipline. Some of them are fairly popular accounts but they are all sufficiently scholarly that it is possible to track down their sources and confront them with alternative perspectives. None of them are by popularisers in the vein of Gladwell or Pinker.
### ‘Against the Grain: A Deep History of the Earliest States’ by James C Scott, 2018
Scott is best known for ‘Seeing Like a State’ but this is a much more important and in many ways better book. His main thesis is ‘Everything we thought about the invention of agriculture and its role in the formation of early civilisations is wrong.’
In this book, Scott summarises recent decades of research on the emergence of agriculture and emergence of early states and finds that we cannot trust any of our assumptions. The early states were temporary, partial and patchy. They cannot be seen as a final stage in some sort of a process of social evolution. A point elaborated by Yoffee below in greater detail.
My main impression from this book is how recent the dominance of state control is. Until about 1500, most people lived outside the control of the great civilisational behemoths. And this was, for many of them, a conscious choice. As Scott described in his earlier book ‘The Art of Not Being Governed’ (also well worth a read).
Similarly to Diamond, Scott also focuses on the importance of certain crops but from the perspective of their utility for taxation. This point is elaborated in Graeber’s ‘Debt’ (see below).
You can see Scott speak about many of these points in several lectures.
- How the Grain Domesticated Us (https://www.youtube.com/watch?v=QO__r8Q0bmU)
- Beyond the Pale: The Earliest Agrarian States and “their Barbarians” (https://www.youtube.com/watch?v=u2ukte-je8k)
- A Short Account of the Deep History of State Evasion (https://www.youtube.com/watch?v=vb8qQcShWUk)
- On the Art of Not Being Governed (https://www.youtube.com/watch?v=RNkkEU7EoOk)
Note: I wrote a review of this book for the Czech daily Lidové noviny.
#### Illustrative quotes
“Contrary to earlier assumptions, hunters and gatherers—even today in the marginal refugia they inhabit—are nothing like the famished, one-day-away-from-starvation desperados of folklore. Hunters and gathers have, in fact, never looked so good—in terms of their diet, their health, and their leisure. Agriculturalists, on the contrary, have never looked so bad—in terms of their diet, their health, and their leisure.”
“In unreflective use, “collapse” denotes the civilizational tragedy of a great early kingdom being brought low, along with its cultural achievements. We should pause before adopting this usage. Many kingdoms were, in fact, confederations of smaller settlements, and “collapse” might mean no more than that they have, once again, fragmented into their constituent parts, perhaps to reassemble later. In the case of reduced rainfall and crop yields, “collapse” might mean a fairly routine dispersal to deal with periodic climate variation. Even in the case of, say, flight or rebellion against taxes, corvée labor, or conscription, might we not celebrate—or at least not deplore—the destruction of an oppressive social order?”
“until the past four hundred years, one-third of the globe was still occupied by hunter-gatherers, shifting cultivators, pastoralists, and independent horticulturalists, while states, being essentially agrarian, were confined largely to that small portion of the globe suitable for cultivation. Much of the world’s population might never have met that hallmark of the state: a tax collector.”
“Where grain, and therefore agrarian taxes, stopped, there too did the state’s power begin to degrade. The power of the early Chinese states was confined to the arable drainage basins of the Yellow and Yangzi Rivers. […] The territory of the Roman Empire, for all its imperial ambitions, did not extend much beyond the grain line.”
much that passes as collapse as, rather, a disassembly of larger but more fragile political units into their smaller and often more stable components.
### ‘Seven Myths of the Spanish Conquest’ by Matthew Restall, 2003
The title of the book says it all. Almost anything we say (and Jared Diamond said) about the likes of Columbus, Cortez or Pizarro is wrong. Factually and structurally. Perhaps the most important myth Restall presents is that of ‘completion’. The Spanish and later other conquests were more a case of expanding enclaves and negotiations. To imagine the conquistadors as ruling a geographic area in the same way a modern state governs its territory is completely misleading. This is also a point repeated in Yoffee and Scott with respect to ‘ancient civilisations’.
The other point made by Restall is the complete dependence of the European invaders on local political aliances and the relative ineffectiveness and ultimate irrelevance of their ‘technology’. We see this expanded in Thornton and Sherman into other contexts
You can hear Restall talk about many of the same themes in an interview about his more recent book “When Montesuma met Cortez” in this New Books Podcast (https://newbooksnetwork.com/matthew-restall-when-montezuma-met-cortes-the-true-story-of-the-meeting-that-changed-history-ecco-2018/).
There is also an illustrated lecture available on YouTube (https://www.youtube.com/watch?v=KhHJituiAcQ) that covers the same topics.
#### Illustrative quotes
“Looking at Spanish America in its entirety, the Conquest as a series of armed expeditions and military actions against Native Americans never ended.”
“Only very gradually did community autonomy erode under demographic and political pressures from non-native populations. From the native perspective, therefore, the Conquest was not a dramatic singular event, symbolized by any one incident or moment, as it was for Spaniards. Rather, the Spanish invasion and colonial rule were part of a larger, protracted process of negotiation and accommodation.”
### ‘Empires of the Weak: The Real Story of European Expansion and the Creation of the New World Order’ by Jason Sharman, 2019
The central thesis here is that the relationship between the European conquerors and the conquered around the world was very different from the traditional stories. It was not sudden overwhelming military force but gradual exploitation of local political conditions taking place over the course of centuries that resulted in the world we see today.
Most importantly, the thesis of political competition in Europe resulting in European dominance by 1800 purely through superiority of Western military technology is completely dismantled. European weapons made little difference until the 1800s. Updated based on Reddit comments. (https://www.reddit.com/r/slatestarcodex/comments/d1gbsj/so_you_think_you_have_a_historical_analogy/)
Sharman is a political scientist, so perhaps could be accused of moonlighting outside his core expertise, but we’ll see that this thesis is repeated again and again in many of the other books on this list from various perspectives.
I could not find any videos or audio recordings of Sharman about the book. But I’m sure some will appear, soon.
#### Key quote
‘Europeans did not enjoy any significant military superiority vis-à-vis non-Western opponents in the early modern era, even in Europe. Expansion was as much a story of European deference and subordination as one of dominance. Rather than state armies or navies, the vanguards of expansion were small bands of adventurers or chartered companies, who relied on the cultivation of local allies.’
### ‘The Silk Roads: A New History of the World’ by Peter Frankopan, 2015
‘The Silk Roads’ is the history of the world that should be the core textbook for anyone interested in the balance of events. It provides the same correction to the shape of history that an alternative projection gives to the distortions taught to us by the Mercator of school atlases (https://geoawesomeness.com/best-map-projection).
Frankopan’s book on the First Crusade is also extremely eye-opening and worth a read. It is the one that most balances the perspectives of east, west and the Byzantines.
Here are some places where you can see Frankopan talk about his book:
- The Silk Roads: Questioning the Eurocentric view of history (https://www.youtube.com/watch?v=NJ54ojX5zlM)
- Similar talk given at Yale (https://www.youtube.com/watch?v=NJ54ojX5zlM)
- Discussion of the book at the OU (https://podcasts.ox.ac.uk/silk-roads-new-history-world).
Jerry Brotton’s ‘This Orient Isle’ could be thought of as a companion book in that it rethinks the position of Britain in this newly rebalanced history. You can watch Brotton talk about his 2016 book ‘in an online lecture (https://www.youtube.com/watch?v=0tW-tvK9t2s). Brotton’s ‘History of the World in 12 Maps’ also adds new perspectives on the orientation of the world.
#### Illustrative quotes
We think of globalisation as a uniquely modern phenomenon; yet 2,000 years ago too, it was a fact of life, one that presented opportunities, created problems and prompted technological advance.
Rome’s transition into an empire had little to do with Europe or with establishing control across a continent that was poorly supplied with the kind of resources and cities that were honeypots of consumers and taxpayers. What propelled Rome into a new era was its reorientation towards the Eastern Mediterranean and beyond. Rome’s success and its glory stemmed from its seizure of Egypt in the first instance, and then from setting its anchor in the east – in Asia.
the ancient world was much more sophisticated and interlinked than we sometimes like to think. Seeing Rome as the progenitor of western Europe overlooks the fact that it consistently looked to and in many ways was shaped by influences from the east.
Cities like Merv, Gundesāpūr and even Kashgar, the oasis town that was the entry point to China, had archbishops long before Canterbury did. These were major Christian centres many centuries before the first missionaries reached Poland or Scandinavia.
Baghdad is closer to Jerusalem than to Athens, while Teheran is nearer the Holy Land than Rome, and Samarkand is closer to it than Paris and London.
### ‘Lost Enlightenment: Central Asia’s Golden Age from the Arab Conquest to Tamerlane’ by F. Frederick Starr, 2013
Starr’s book was a real revelation. I had spent a lot of time in Central Asia and read some history of the region. But other than Samarkand, all of that history has now been lost. And I didn’t get a sense that the people living in the region knew much about it.
Much like Davies in ‘Vanished Kingdoms’ in Europe, Starr shows on a global scale how even major civilisations with real impact can disappear without much trace. But even more importantly, he shows that the trajectory of ‘modern’ intellectual development was much more complex than most people believe.
You can see Starr talk about his book in this online lecture (https://www.youtube.com/watch?v=qSDqjUoH67M).
#### Illustrative quotes
“This was truly an Age of Enlightenment, several centuries of cultural flowering during which Central Asia was the intellectual hub of the world. India, China, the Middle East, and Europe all boasted rich traditions in the realm of ideas, but during the four or five centuries around AD 1000 it was Central Asia, the one world region that touched all these other centers, that surged to the fore. It bridged time as well as geography, in the process becoming the great link between antiquity and the modern world.”
“every major Central Asian city at the time boasted one or more libraries, some of them governmental and others private.”
“Above all, Central Asia was a land of cities. Long before the Arab invasion, the most renowned Greek geographer, Strabo, writing in the first century BC, described the Central Asian heartland as ‘a land of 1,000 cities.’”
“At the Merv oasis the outermost rampart ran for more than 155 miles, three times the length of Hadrian’s Wall separating England from Scotland. At least ten days would have been required to cover this distance on camelback.”
### ‘Genghiz Khan and the Making of the Modern World’ by Jack Weatherford, 2004
Jack Weatherford’s portrayal of the Mongol conquests is definitely not non-partisan. He’s with the Mongols. Nevertheless, he opens important vistas about the foundations of modern interconnectedness. This is a good complement to Starr’s covering of the preceding period in the same region.
Here’s a video lecture by Weatherford (https://youtu.be/v81_hm8T92c) about some aspects of this story.
#### Illustrative quotes
In twenty-five years, the Mongol army subjugated more lands and people than the Romans had conquered in four hundred years. Genghis Khan, together with his sons and grandsons, conquered the most densely populated civilizations of the thirteenth century. Whether measured by the total number of people defeated, the sum of the countries annexed, or by the total area occupied, Genghis Khan conquered more than twice as much as any other man in history.
The majority of people today live in countries conquered by the Mongols; on the modern map, Genghis Kahn’s conquests include thirty countries with well over 3 billion people.
Genghis Khan’s empire connected and amalgamated the many civilizations around him into a new world order. At the time of his birth in 1162, the Old World consisted of a series of regional civilizations each of which could claim virtually no knowledge of any civilization beyond its closest neighbor. No one in China had heard of Europe, and no one in Europe had heard of China, and, so far as is known, no person had made the journey from one to the other. By the time of his death in 1227, he had connected them with diplomatic and commercial contacts that still remain unbroken.
### ‘Lies my Teacher Told Me: Everything Your American History Textbook Got Wrong’ by James W. Loewen, 1995
This book is slightly outside the scope of this list, but I thought it would be of interest to those who were educated in the American school system. But many of its points apply to all school history books. It will open your eyes to how little you can trust to what you learned in school and what was then reinforced through popular cultural reflection of history.
Here’s an extended interview with the author (https://www.youtube.com/watch?v=JTTA_41UGUA).
#### Illustrative quotes
“Many history textbooks list up-to-the-minute secondary sources in their bibliographies, yet the narratives remain totally traditional unaffected by recent research.”
“Most Americans tend automatically to equate educated with informed or tolerant. Traditional purveyors of social studies and American history seize upon precisely this belief to rationalize their enterprise, claiming that history courses lead to a more enlightened citizenry. The Vietnam exercise suggests the opposite is more likely true.”
## Comprehensive and/or less accessible
These books require more serious commitment and possibly some comfort with reading relatively dense historical and ethnographic accounts. They are not necessarily poorly written or full of jargon but they are not primarily aimed at an audience too far outside the profession of the author (except ‘Debt’ which I included here because it is so long).
### ‘A Cultural History of the Atlantic World: 1250 - 1820’ by John K. Thornton, 2012
This is a truly impressive historical synthesis that covers an extensive geographic area as well as a significant stretch of time. It provides detailed elaborations of the central thesis of Sharman’s and Restall’s books and should be consulted every time we feel like we want to make a general statement about the developments in that region and in that time. Which we do all the time.
A podcast interview about this book (https://newbooksnetwork.com/john-k-thornton-a-cultural-history-of-the-atlantic-world-1250-1820-cambridge-up-2012-3/) from the New Books Network will give a good sense of what the book is about.
You can also hear Thornton speak on a related topic in this YouTube lecture on the Slave trade (https://www.youtube.com/watch?v=kxUjIt2EmRA).
#### Illustrative quotes
“Europeans did not possess decisive advantages over any of the people they met, even though their sailing craft were indeed capable of nautical achievements that no other culture up to that time was able to perform.”
“there was really no economic Third World at the time of European expansion in the fifteenth and sixteenth century, if one uses proxy measures of average quality of life as a guide. The crucial quality-of-life determinant was, in fact, social and economic stratification.”
“African states had the upper hand if the game of force was to be played. Although Europeans often fortified their “factories,” as trading posts were usually called, these fortifications could not resist an attack by determined African authorities.”
“Slow-firing weapons cannot allow small numbers of people to defeat larger numbers unless other factors are in play.”
“Cavalry are most effective only when massed in sufficient numbers to inflict sustained casualties on fleeing infantry, and the dozens and on occasion low hundreds of mounted men in Spanish service did not meet this decisive threshold. Native Americans were reasonably quick in establishing tactical countermeasures against the horsemen after the initial encounters.”
### ‘Myths of the Archaic State: Evolution of the Earliest Cities, States, and Civilizations’ by Norman Yoffee, 2005
This was perhaps the most embarrassingly eye-opening book for me given that I started out my early adult life by studying Egyptology. Yoffee, building on his work and that of others, shows the limits of what a so-called ‘ancient civilisation’ was and could have been. Collapses and interregna were all much less of tragedies for all involved - point made by Scott in a more accessible way. His reimagining of the position of Hammurabi as a political and literary rather than a legal document was just one of the many myths that this book burst for me.
Norman Yoffee speaks about new perspectives on the collapse (https://www.youtube.com/watch?v=R0TvJx1F7QM) summarizing his more recent work.
#### Illustrative quotes
“[Myth of the archaic states include:] (1) the earliest states were basically all the same kind of thing (whereas bands, tribes, and chiefdoms all varied within their types considerably);(2) ancient states were totalitarian regimes, ruled by despots who monopolized the flow of goods, services, and information and imposed “true” law and order on their powerless citizens; (3) the earliest states enclosed large regions and were territorially integrated; (4) typologies should and can be devised in order to measure societies in a ladder of progressiveness; (5) prehistoric representatives of these social types can be correlated, by analogy, with modern societies reported by ethnographers; and (6) structural changes in political and economic systems were the engines for, and are hence necessary and sufficient conditions that explain, the evolution of the earliest states.”
That the laws of Hammurabi were copied in Mesopotamian schools for over a millennium after Hammurabi’s death attests to the literary success of the composition and has nothing to do with its juridical applicability. […] There is no mention of the code of Hammurabi in the thousands of legal documents that date to his reign and those of his immediate successors.
Order could not survive the frequent shocks it suffered if people were not able to construct the institutions of legitimacy and to determine the quality of illegitimacy. Legitimacy normally invokes the past as something that is absolute and that acts as a point of reference for the present, normally by transmuting the past into some form of the present.
### ‘Debt: The First Five Thousand Years’ by David Graeber, 2011
I think this is perhaps the best intro to modern anthropological thinking in general. It is very readable and accessible but also very comprehensive. It certainly has its agenda but Graeber tells a convincing story that undermines the classical thinking about the role of exchange in maintaining civilisations. It is easy to get bogged down in the discussions about the nature of money when discussing this book but what it really does is show the great variety of ways in which people relate to each other.
Graeber gave a lecture on his book at Google which is available on YouTube (https://www.youtube.com/watch?v=R0TvJx1F7QM). But this book works best when read as a whole.
#### Illustrative quotes
there is good reason to believe that barter is not a particularly ancient phenomenon at all, but has only really become widespread in modern times. Certainly in most of the cases we know about, it takes place between people who are familiar with the use of money, but for one reason or another, don’t have a lot of it around.
Through most of history, when overt political conflict between classes did appear, it took the form of pleas for debt cancellation—the freeing of those in bondage, and usually, a more just reallocation of the land.
“Kingdoms rise and fall; they also strengthen and weaken; governments may make their presence known in people’s lives quite sporadically, and many people in history were never entirely clear whose government they were actually in. … It’s only the modern state, with its elaborate border controls and social policies, that enables us to imagine “society” in this way, as a single bounded entity.”
there are three main moral principles on which economic relations can be founded, all of which occur in any human society, and which I will call communism, hierarchy, and exchange.
“communism” is not some magical utopia, and neither does it have anything to do with ownership of the means of production. It is something that exists right now—that exists, to some degree, in any human society, although there has never been one in which everything has been organized in that way, and it would be difficult to imagine how there could be. All of us act like communists a good deal of the time. None of us acts like a communist consistently.
“baseline communism”: the understanding that, unless people consider themselves enemies, if the need is considered great enough, or the cost considered reasonable enough, the principle of “from each according to their abilities, to each according to their needs” will be assumed to apply.
In many periods—from imperial Rome to medieval China—probably the most important relationships, at least in towns and cities, were those of patronage.
### ‘Vanished Kingdoms: The Rise and Fall of Nations’ by Norman Davies, 2010
Most of the books on the list focus on forgotten, misunderstood or ignored aspects of global history or culture. But Davies shows that even in our own backyard, entire kingdoms vanished without a trace in our consciousness. Perhaps, I should have chosen his ‘Europe: A History’ which is also revisionist in that it places Europe’s cultural, geographic and historical center of gravity much further east and south than is typical. But I found this book much more revelatory and impactful for the purposes of this list.
Davies gave a lecture about his book at the LSE (https://www.mixcloud.com/lse/europes-vanished-kingdoms-audio/)which is available as a recording.
A brief interview with Davies about this book (https://www.youtube.com/watch?v=C95Yb2A27BY) is available on YouTube.
#### Illustrative quotes
“As soon as great powers arise, whether the United States in the twentieth century or China in the twenty-first, the call goes out for offerings on American History or Chinese History, and siren voices sing that today’s important countries are also those whose past is most deserving of examination, that a more comprehensive spectrum of historical knowledge can be safely ignored.”
Most importantly, students of history need to be constantly reminded of the transience of power, for transience is one of the fundamental characteristics both of the human condition and of the political order.
Popular memory-making plays many tricks. One of them may be called ‘the foreshortening of time’. Peering back into the past, contemporary Europeans see modern history in the foreground, medieval history in the middle distance, and the post-Roman twilight as a faint strip along the far horizon.
One has to put aside the popular notion that language and culture are endlessly passed on from generation to generation, rather as if ‘Scottishness’ or ‘Englishness’ were essential constituents of some national genetic code.
To all who have been seduced by the concept of ‘Western Civilization’, therefore, the Byzantine Empire appears as the antithesis – the butt, the scapegoat, the pariah, the undesirable ‘other’.6 Although it formed part of a story that lasted longer than any other kingdom or empire in Europe’s past, and contains in its record a full panoply of all the virtues, vices and banalities that the centuries can muster, it has been subjected in modern times to a campaign of denigration of unparalleled virulence and duration.
### ‘The Inheritance of Rome: A History of Europe from 400 to 1000’ by Chris Wickham, 2009
The ‘Fall of Rome’ and the subsequent ‘dark ages’ have been one of the big obsessions of historical introspection for centuries. They are the frequent source domain of civilizational analogies even though, as Chris Wickham shows, almost nothing we think of as a given holds up. This book is just one of many in recent historical scholarship that revisits the notion of the dark ages and shines a light on the period of ‘late Rome’ as seen from the perspective of its own time. Of course, many controversies remain but the change in emphasis seems incontrovertible. There are echoes of similar points in the early chapters of Davies’ ‘Vanished Kingdoms’.
I could not find any lectures on the subject by Wickham but some of these questions were raised in this panel he chaired on the middle ages (http://dyoutube.com/watch?v=kAz8w2IKHHM).
Note: I wrote a review of this book for the Czech daily Lidové noviny.
There are also a number of lecture series on this period that reflect the latest scholarship in The Great Courses from the Teaching Company that are available via Audible, as well. I particularly recommend those by Kenneth Harl (https://www.thegreatcourses.com/search/?q=kenneth+harl).
#### Illustrative quotes
“Anyone in 1000 looking for future industrialization would have put bets on the economy of Egypt, not of the Rhineland and Low Countries, and that of Lancashire would have seemed like a joke.”
Byzantine ‘national identity’ has not been much considered by historians, for that empire was the ancestor of no modern nation state, but it is arguable that it was the most developed in Europe at the end of our period.
the East remained politically and fiscally strong, and eastern Mediterranean commerce was as active in 600 as in 400.
Far from ‘corruption’ being an element of Roman weakness, this vast network of favours was one of the main elements that made the empire work. It was when patronage failed that there was trouble.
The Persian state was almost as large as the Roman empire, extending eastwards into central Asia and what is now Afghanistan; it is much less well documented than the Roman empire, but it, too, was held together by a complex tax system, although it had a powerful military aristocracy as well, unlike Rome.
### ‘The Anthropology of Eastern Religions: Ideas, Organizations, and Constituencies’ by Murray Leaf, 2014
This was a late addition to this list and it is an imperfect volume in that, as one reviewer put it (https://www.tandfonline.com/doi/full/10.1080/10477845.2015.1071584): “[its] worthwhile aims are met unevenly, resulting in a book that is certainly informed and informative, but often inconsistent in tone and level of analysis.” But I think its core message in chapter one of religion as a social institution which has much more in common with others than traditional religious studies would have us believe.
I couldn’t find any interviews or lectures. But I believe that this interview with Russell McCutcheon (https://newbooksnetwork.com/william-arnal-and-russell-t-mccutcheon-the-sacred-is-the-profane-the-political-nature-of-religion-oxford-up-2013-3/) about the limits of religious studies would provide a useful complement.
#### Illustrative quotes
The world religions are complex social phenomena. They use ideas of several different kinds. They include substantial systems of physical infrastructure. They have provisions for economic support. They embody their own systems of scholarship. They produce propaganda and they are politically important in many different ways. From time to time their leaders in various places have commanded armies and conducted wars. This cannot be explained simply by reviewing them as so many sets of beliefs.
The most conspicuous problem in contemporary comparative religion is that they underrate diversity. […] One result is to overstate what the major religions have in common with each other while understating or ignoring what they have in common with traditions considered non-religious.
The general class of cultural phenomena to which world religions belong can be described as large-scale, translocal, multi-organizational, professionalized cultural complexes.
Virtually all Japanese have recourse to the ideas and organizations of Buddhism and Shinto, and for the most part this is also true of Confucianism. […] Japanese parks are Shinto and the system of Shinto shrines in Japan has much the same place in Japanese emotional life as the system of national parks does for Americans. Confucian ideas are important in administrative and professional contexts.
### ‘Europe and the People without History’ by Eric R Wolf, 1982
This is the oldest book on the list and it has inspired many others.
Unlike Graeber, reading Wolf is hard going. This is certainly for the committed but it repays the effort. Even just looking at the maps showing the intricate trade routes going from the heart of Africa to the Baltic Sea is eye-opening.
Wolf’s central point is also the central point all the authors on this list return to again and again. We invented history based on the things that were easy to see. But this was very much looking for the keys under the lamppost where the light was and not where we lost them. Wolf (similarly to Graeber and Scott) has a distinctly untraditional politics leaning to the left (if perhaps not as much to anarchism).
#### Illustrative quotes
“Africa south of the Sahara was not the isolated, backward area of European imagination, but an integral part of a web of relations that connected forest cultivators and miners with savanna and desert traders and with the merchants and rulers of the North African settled belt. This web of relations had a warp of gold, “the golden trade of the Moors,” but a weft of exchanges in other products. The trade had direct political consequences. What happened in Nigerian Benin or Hausa Kano had repercussions in Tunis and Rabat. When the Europeans would enter West Africa from the coast, they would be setting foot in a country already dense with towns and settlements, and caught up in networks of exchange that far transcended the narrow enclaves of the European emporia on the coast. We can see such repercussions at the northern terminus of the trade routes in Morocco and Algeria. Here one elite after another came to the fore, each one dependent on interaction with the Sahara and the forest zone. Each successive elite was anchored in a kin-organized confederacy, usually mobilized around a religious ideology.”
## Pre-cursors and proto-revisionists
In many ways, almost any history is revisionist history. Each generation writes its own history books to reflect new knowledge but also new perspectives. Most history book authors feel they have something new with which to contribute and that can revise current understanding of the subject matter.
So it is not surprising that even revisionism in the vein that I’m looking at here is not just a matter of the last 20 or so years.
Much of the current revision was inspired by‘The Great Transformation’ by Karl Polanyi (https://en.wikipedia.org/wiki/The_Great_Transformation_(book)) published in 1944 which in turn rests on many of the anthropological revisions started by people like Franz Boas (https://en.wikipedia.org/wiki/Franz_Boas) in the US and Bronisław Malinowski (https://en.wikipedia.org/wiki/Bronisław_Malinowski).
There is a continued thread of back and forth since at least then. Marshall Sahlins’ ‘Stone Age Economics’ from 1972 (with papers going back to the mid 1960s) started much revision and revision of the hunter gatherer condition. And so on.
At the same time but independently, people like Joseph Needham (https://en.wikipedia.org/wiki/Joseph_Needham) were painstakingly collecting data on the great civilizations of the ‘East’ which can now give us a fuller and more balanced picture of the world.
We should also not forget the work that has gone into revising the simplistic view of the 19th and 20th centuries, most notably to do with the emergence of the nation state. Here names such as Eric Hobsbawn, Ernest Gellner, Miroslav Hroch and, of course, Benedict Anderson, come to mind. And then there are the many people who are rethinking more recent events such as Timothy Snyder or Antony Beevor.
The list just goes on.
## The other side of the coin
Of course, there is also the other side. Historical revisionists with grand schemes and overarching historical narratives. I’ve already mentioned Jared Diamond but also worth reading is Ian Morris. I’ve critiqued some of their work in my thesis proposing the metaphor of ‘History as Weather’ (https://medium.com/metaphor-hacker/history-as-weather-a-fractal-theory-of-history-for-ian-morris-jared-diamond-and-cgp-grey-45b5503486c5). There are also people like Niall Ferguson (https://en.wikipedia.org/wiki/Niall_Ferguson) who cannot be doing with all this rebalancing and want to put ‘the West’ back at the centre of things. I found his attempt in Civilization extremely unconvincing (/2011/06/killer-app-is-a-bad-metaphor-for-historical-trends-good-for-pseudoteaching) but he is a prominent voice in the anti-revisionist camp.
Note: I wrote a joint review of Morris and Ferguson for the Czech daily Lidové noviny under the title of ‘New historical eschatology’. I also wrote a positive review of Jared Diamond’s ‘Collapse’ soon after it came out which I would now revise significantly - see Questioning Collapse (https://www.amazon.co.uk/Questioning-Collapse-Resilience-Ecological-Vulnerability/dp/0521733669) and After Collapse (https://www.amazon.co.uk/After-Collapse-Regeneration-Complex-Societies/dp/0816529361).
Steven Pinker’s ‘Better Angels of Our Nature (https://en.wikipedia.org/wiki/The_Better_Angels_of_Our_Nature)‘ is another example of a grand sweep of history that tries to make the progress of humanity appear more directional and straightforward than the books on this list suggest it is or can be. And we should not forget the great systematisers of the 1990s Fukuyama (https://en.wikipedia.org/wiki/The_End_of_History_and_the_Last_Man) andHuntington (https://en.wikipedia.org/wiki/Clash_of_Civilizations).
The temptation to discover the key to what makes human history tick is great. FromToynbee (https://en.wikipedia.org/wiki/Arnold_J._Toynbee#Academic_and_cultural_influence) to Hari Seldon (https://en.wikipedia.org/wiki/Hari_Seldon). And it is not difficult to discover interesting patterns (https://slatestarcodex.com/2019/08/12/book-review-secular-cycles).
The picture that the books on this list paint is that grand narratives of history do not stand up well to scrutiny. They may provide a useful lens through which to view the past, or more often the present. But there is always another grand narrative just around the corner.
In thinking about the predictive utility of history (https://medium.com/metaphor-hacker/history-as-weather-a-fractal-theory-of-history-for-ian-morris-jared-diamond-and-cgp-grey-45b5503486c5), I asked: “So what is the point of history then? Its accurate predictions are not very useful and its useful predictions are not very accurate.” History and ethnography show us the range of possible ways of being human. They don’t tell us what to do next or how to be, but they are essential components of our never-ending quest to find out what we could be.
## Turing tests in Chinese rooms: What does it mean for AI to outperform humans
Date: 2019-07-23
URL: https://metaphorhacker.net/2019/07/turing-tests-in-chinese-rooms-what-does-it-mean-for-ai-to-outperform-humans/
Categories: AI, Linguistics, Extended writing
Tags: featured
## TLDR;
- Reports that AI beat humans on certain benchmarks or very specialised tasks don’t mean that AI is actually better at those tasks than any individual human.
- They certainly don’t mean that AI is approaching the task with any of the same understanding of the world people do.
- People actually perform 100% on the tasks when administered individually under ideal conditions (no distraction, typical cognitive development, enough time, etc.) They will start making errors only if we give them too many tasks in too short a time.
- This means that just adding more of these results will NOT cumulatively approach general human cognition.
- But it may mean that AI can replace people on certain tasks that were previously mistakenly thought to require general human intelligence.
- All tests of artificial intelligence suffer from Goodart’s law.
- A test more closely resembling an internship or an apprenticeship than a gameshow may be a more effective version of the Imitation Game.
- Worries about ‘superintelligence’ are very likely to be irrelevant because they are based on an unproven notion of arbitrary scalability of intelligence and ignore limits on computability.
## Reports of my intelligence have been greatly exaggerated
Over the last few years, there have been various pronouncements about AI being better than humans at various tasks such as image recognition, speech transcription, or even translation. And that’s not even taking into account bogus winners of the Turing test challenge. To make things worse, there’s always the implication that this is means machine learning is getting closer to human learning and artificial intelligence is only a step away from going general.
All of those reports were false. Every. Single. One. How do we know this? Well, because none of them were followed by “and therefore we have decided to replace all humans doing job X with machine learning algorithms”. But even if this were the case, it still would not necessarily mean that the algorithm outperformers humans at the task. Just that it can outperform them at the task when it is repeated time after time and the algorithm ends up making fewer mistakes because, unlike people, it does not get tired, distracted, or simply ticks the wrong box.
But even if the aggregate number of errors is lower for a machine learning algorithm, it may still not make sense to use it because it makes qualitatively different errors. Errors that are more random and unpredictable are worse than more systematic errors that can be corrected for. Also, because AI has no metacognitive mechanisms to identify its errors by doing a ‘sense check’. This often makes correcting AI-generated transcripts difficult to correct because it makes errors that don’t make intuitive sense.
## Pattern matching in radiology and law
The closest machine learning has gotten to outperforming humans doing real jobs is in radiology. (I’m discounting games like Go, here.) But even here it only equalled the performance of the best experts. However, this could easily be enough. But interpreting X-Rays is an extremely specialised task that requires lots of training and has a built-in error rate. It is a pattern recognition exercise, not a general reasoning exercise. All the general reasoning about the results of the X Rays still has to be delegated to the human physician.
In a similar instance, AI could notice inconsistencies in complex contracts better than lawyers. Again, this is very plausible, but again this was a pattern-matching exercise with a machine pitted against human distractability and stamina. Definitely impressive, useful, and not something expected even a few years ago. But not in any meaningful ways replacing the lawyer any more than a form to draw up a contract I downloaded from the internet does.
This is definitely a case where an AI can significantly augment what an unassisted human can do. And while it will not replace radiologists or lawyers as a category, it could certainly greatly decrease their numbers.
## Machine learning to the test
So on very specialised tasks involving complex pattern recognition, we could say that AI can genuinely outperform humans.
But in all the instances involving language and reasoning tasks, even if an AI beats humans on a test, it does not actually ‘outperform’ them on the task. That’s because tests are always imperfect proxies for the competence they measure.
For example, native speakers often don’t get 100% on English proficiency tests and can even do worse than non-native speakers in certain contexts. Why? Three reasons: 1. They can imagine contexts not expected of non-native speakers. 2. The non-native speakers have been practicing taking these tests a lot so they make fewer formal mistakes.
We are facing exactly the same problems when comparing machine learning and human performance based on tests designed to evaluate machine learning. Humans are the native speakers and they perform 100% on all the tasks in their daily lives. But their performance seems less than perfect in test conditions.
### BLEU and overblown claims about Machine Translation
Sometimes the problem is with a poorly designed test. This is the case with the common measure of machine translation called BLEU (Bi-Lingual Evaluation Understudy) (https://en.wikipedia.org/wiki/BLEU). BLEU essentially measures how many similar words or word pairs there are in the translation by machine when compared to a reference corpus of human translations. It is obvious that this is not a good metric of quality of translation. It can easily assign a lower score to a good translation and a high score to a patently bad one. For instance, it would not notice that the translation missed a ‘not’ and gave the opposite meaning.
What human translators do is translate whole texts NOT sentences. This sometimes means they drop things, add things, rearrange things. This involves a lot of judgment and therefore no two translations are ever the same. And outside trivial cases they’re never perfect. But a reliable translator can make sure they convey the key message and they could provide footnotes to explain where this was not possible. Machine learning can get surprisingly good at translating texts by brute force. But it is NOT reliable because it operates with no underlying understanding of the overall meaning of the text.
That’s why we can easily dismiss Microsoft’s claim that their English-to-Chinese interpreter outperformed human translators. That is only because they used the BLEU metric to make this claim rather than professional translators evaluating the quality of AI output against that of other professional translators on any test. And since Microsoft has yet to announce that it is no longer using human interpreters when its executives visit China, we can safely assume that this ‘outperform’ is not real.
Now, could a machine translation ever get good enough to replace human translators? Possibly. But it is still very far from that for texts of any complexity. Transformers are very promising at improving the quality of the translation but they still only match patterns. To translate you need to make quite rich inferences and we’re nowhere near this.
### GLUE and machine understanding come unstuck
Speaking of inferences. How good is AI at making those? Awful. Here we have another metric to look at: GLUE! Unlike BLEU which is a really bad representation of the quality of translation, GLUE (General Language Understanding Evaluation) (https://gluebenchmark.com) is a really good representation of human intelligence. If you wanted to know what are the components of human intelligence, you could do a lot worse than look at the GLUE test.
But the GLUE leaderboard (https:gluebenchmark.com/submission/xfBamAUrEBY0BWIVLx2uuGSmKvI2/-LXoXWPvqioRKAxCjt5W) and it comes 4th with 87.1% score. This puts it 1.4% behind the leader which is Facebook at 88.5%. So, it’s done. AI has not only reached human level of reasoning, it has surpassed them! Of course, not. Apart from the fact that we don’t know how much of a difference in reasoning ability 1% is, this tells us nothing about human ability to reason when compared to that of a machine learning model. Here’s why.
## How people and machines make errors
I would argue that a successful machine learning algorithm does not actually outperform humans on these tasks even if it got 100%. Because humans also get 100% but they also devised the test.
Isn’t this a contradiction? How can humans get 100% if they consistently score in the mid-80s when given the test. Well, humans designed the test and the correctness criteria. And a machine learning algorithm must match the best human on every single answer to equal them. The benchmark here is just an average of many people over many answers and does not just reflect the human ability to reason but also the human ability to take tests.
Let’s explain by comparing what it means when a human makes an error on a test and when a machine does. There are three sources of human error: 1. Erroneous choice when knowing the right answer (ie clicking a when meaning to click b), 2. Lack of attention (ie choosing a because we didn’t spend enough time reading the task to choose correctly), 3. Overinterpretation (providing context in our head that makes the incorrect answer make sense).
These benchmarks are not Mensa tests, they measure what all people with typical linguistic and cognitive development can do. Let’s take the Windograd Schema test as an example. Here’s an often-quoted example:
The trophy didn’t fit into the suitcase because itwas too big.
The trophy didn’t fit into the suitcase because itwas too small.
It is very possible that out of 100 people, 5 would get this wrong because they click the wrong answer, 10 because they didn’t process the sentence structure correctly and 1 because they constructed a scenario in their head in which it is normal for suitcases to be smaller than the thing in them (as in Terry Pratchett’s books).
But not a single one got it wrong because they thought that a thing can be bigger than the thing it fits in.
Now, when a machine learning model gets it wrong, it does it because it miscalculated a probability based on an opaque feature set it constructs from lots of examples. When you get 2 people together, they can always figure out the right answer and discuss why they did it wrong. No machine learning algorithm can do that.
This becomes even more obvious when we take an example from the actual GLUE benchmark:
Maude and Dora had seen the trains rushing across the prairie, with long, rolling puffs of black smoke streaming back from the engine. Their roars and their wild, clear whistles could be heard from far away. Horses ran away when they came in sight.
So what does the ‘they’ refer to here? The obvious candidate here is ‘trains’. But it is easy to imagine that a person could click the option where ‘puffs of black smoke’ or even ‘Maude and Dora’ are the antecedent. That’s because both of those can be ‘seen’ and could theoretically cause horses to run away. If this is the 10th sentence I’m parsing in a go, I may easily shortcut the rather complex syntactic processing. I can even see someone choosing “whistles” even though they cannot “come in sight” but are a very strong candidate for causing horses to run away. But nobody would choose ‘horses’ unless they misclicked. A machine learning algorithm very easily could do this simply because ‘they’ and ‘horses’ match grammatically.
But all of this is actually irrelevant, because of how the ML algorithms are tested. They are given multiple pairs or sentences and asked to say 1 or 0 on whether they match or not. So some candidate sentences above are “Horses ran away when the trains came in sight.”, “Horses ran away when Maude and Dora came in sight.” or “Horses ran away when the whistles came in sight.” What it does NOT do is ask “Which of the words in the sentence does ‘they’ refer to?” Because the ML model has no understanding of such questions. You would have to train it for that task separately or just write a sequential algorithm to process these questions.
What people running these contests also cannot do is ask the model to explain their choice in a way that would show some understanding. There is a lot of work being done on interpretability, but this just spits out a bunch of parameters that have to be interpreted by people. Game, set and match to humans.
## Chinese room revisited
But let’s also think about what it means for a neural network model to get things right. This brings us back to Searl’s famous Chinese room argument. Every single choice a model makes has assigned a probability and even quite ridiculous choices have a non-zero chance of being right in the model. Let’s look at another common example:
The animal didn’t cross the road because it was too busy.
Here it is sensible to assign it to ‘road’ because it makes the most sense but one could imagine a context in which we could make it refer to ‘the animal’. Animals can be thought of as busy and we can imagine that this could be a reason for not crossing the road. But we know with 100% certainty that it does not refer to ‘the’ or even ‘cross’. Yet, a neural model has no such assurance. It may never choose ‘the’ in practice as the antecedent for ‘it’ but it will never completely discount it, either.
So, even if the model got everything right. We could hardly think of it as making human-like inferences unless it could label certain antecedents as having 0% probability and others (much rarer) as having 100%. (Note: Programming it to change 10% to 0% or 90% to 100% does not count.)
This feels like a very practical expression of Searl’s Chinese room argument (https://en.wikipedia.org/wiki/Chinese_room) albeit in a weak form. Neural networks pose a challenge to Searl because their algorithmic guts are not as exposed as those of the expert systems of Searl’s time. But we can still see echoes of their lack of actual human-like reasoning in their scores.
## Is a test of artificial intelligence possible under Goodhart's Law?
I once attended a conference on AI risk where a skeptic said he wasn’t going to worry “until an AI could do Winograd schemas”. This referred to a test of common sense and linguistic ambiguity that AIs have long been famously bad at. NowMicrosoft claims (https:slatestarcodex.com/2019/06/26/links-6-19))
This post was inspired by the above remark by Scott Alexander. I wanted to explain why even the Winograd challenge being conquered is not enough in and of itself.
AI proponents constantly complain of sceptics’ shifting standards. When AI achieves a benchmark, everybody scrambles to find something else that could be required of it before it gets a pass. And I admit that I may have made a claim similar to that of the AI researcher quoted by Scott Alexander when I was writing about the Winograd schemas.
But the problem here is not that machines became intelligent and everybody is scrambling to deny the reality. The problem is that they got better at passing the test in ways that nobody envisioned when the test was designed. All this while taking no steps towards actual intelligence. Although with a possible increase in practical utility.
This is the essence of Goodhart’s Law: “When a measure becomes a target, it ceases to be a good measure.” The Winograd Schema Challenge seemed so perfect. Yet, I can imagine a machine learning getting good at passing the challenge but still not actually having any of the cognition necessary to really deal with the tasks in real life. In the same way that IBM Watson got really good at Jeopardy but failed at everything else.
None of this is to say that machine learning could not get good enough at performing many tasks that were previously thought to require generalised cognitive capacity. But when machines actually achieve human-level artificial intelligence, we will know. It will not be that hard to tell. But it will not likely happen just because we’re doing more of the same.
The problem with the Turing test or imitation game is not that it cannot produce reliable results on any one run of it. The problem is that if any single test becomes not only the measure but also a target, it is very much possible to focus on passing the test on the surface while bypassing the underlying abilities the test is meant to measure. But the problem is not just with the individual tests but rather in the illusion that we can design a test that will determine AGI level performance simply by reaching an arbitrary threshold.
The current Turing test winners won by misdirection that hid the fact that they refused to answer the questions. This could be fixed by requiring that Grice’s cooperative principle maxims (https://en.wikipedia.org/wiki/Cooperative_principle#Grice's_Maxims) are observed (especially quality and relevance) but even then, I could see a system trained to deal with a single time-bound conversation pass without any underlying intelligence.
As Scott Aaronson showed (https://www.scottaaronson.com/blog/?p=1858), it is possible to defeat a current level AI system simply by asking ‘What is bigger a shoebox or Mount Everest’. But once a pattern of questioning becomes known, it becomes a target and therefore a bad measure.
Similar things happen with all standardised aptitude tests designed so that they cannot be studied for. Job interview techniques designed to get interviewees to reveal their inner strengths and weaknesses. All of these immediately spawn industries of prep schools, instructional guides, etc. That makes them less useful over time (assuming they were all that useful to start with).
## Towards a test by Critical Turing Internship
That’s why the Turing test cannot be a ‘test’ in the traditional sense. At the very least, it cannot be a single test.
History and a lot of human-computer interaction research has also shown that people are very bad at administering the Turing test (or playing the imitation game). But this is paradoxically because they’re very good the very thing the machines have been failing at: meaning making. Because we almost never encounter meaningless symbols but often encounter incomplete ones, we are conditioned to always infer some sort of meaning from any communication. And it is difficult if not impossible to turn it off.
Every time we see a bit of language we automatically imbue it with some meaning. So, any Turing tester must not only be trained in the principles of cognition but also to discard their own linguistic instincts. We don’t know what it will take for a machine to become truly intelligent but we do know that humans are notoriously bad at telling machines apart from other humans. We simply cannot entrust this sort of thing to such feeble foundations.
As I said above, I suspect that by the time machines do achieve human-level performance on these tasks, it will be obvious. We probably won’t need such a test. Assuming we get there which is not a given. But if a test were needed, it could look something like this.
To replace the Turing test, I would like to propose a sort of Turing Internship. We don’t entrust critical tasks in fields like medicine to people who just passed a test but require they prove ourselves in a closely supervised context. In the same way, we should not trust any AI system based on a benchmark.
Any proposed human-level AI system can be placed in multiple real contexts with several well-informed human supervisors who would monitor its performance for a period of weeks or months to allow for any tricks to be exposed. For example, most people after a few weeks with Alexa, Google Assistant or Siri, get a clear picture of its strengths and limitations. Five minutes with Alexa may make you feel like the singularity is here. Five months will firmly convince you that it is nowhere in sight.
But at the moment, we don’t need this. We don’t need months or weeks to evaluate AI for human-level intelligence. We need minutes. I estimate that we would not need to use this kind of AI internship for another 50 years but likely for much much longer. We are too obssessed with the rapid progress of some basic technologies but ignore many examples of stagnation. My favourite here is the Roomba which has been on the market for 17 years now and has hardly progressed at all. Equally, the current NLP technologies have made massive strides in utility but have not progressed towards anything that could be meaningfully described as understanding.
That is not to say that tests like GLUE or even BLUE are completely useless. They can certainly help us compare ML approaches (up to a point). They’re just useless for comparing human performance with those of machine-generated models.
## Note on Nick Bostrom and Superintelligence
One obvious objection to the Turing Internship idea is that if human-level AI is the last step before Bostrom’s ‘Superintelligence’, unleashing it in any real context would be extremely dangerous.
If you believe in this ‘demon in the machine’ option, there’s nothing I can do to convince you. But I personally don’t find Superintelligence in any way persuasive. The reason is that most of the scenarios described are computationally infeasible in the first place. Bostrom does not mention the issue of computability and things like P=NP almost at all. And he completely ignores questions of nonlinear complexity.
It is hard to judge whether a ‘superintelligent’ system could take over the world. But could it predict the weather 20 days out with 1% tolerance of temperature estimates in any location? The answer is most likely not. There may not be enough atoms in the universe to compute the weather arbitrarily precisely more than a few days in advance. Could it predict earthquakes? Could it run an economy more efficiently than an open market relying on price signals? The answers to all those questions are most likely no. Not because the superintelligence is not super enough but because these may not be problems that can be solved by adding ‘more’ intelligence. Assuming that ‘intelligence’ is a linearly scalable property in the first place. It may well be like body size, after a certain amount of increase, it would just collapse onto itself.
Superintelligence requires a conspiracy theorist’s mindset. Not that people who believe are conspiracy theorists. But they assume that complexity can be conquered with intelligence. They don’t believe that humans are ‘smart’ enough to control everything. But they believe that it is inherently possible. Everything we know about complexity, suggests that this is not the case. And that is why I’m not worried.
## Fruit loops and metaphors: Metaphors are not about explaining the abstract through concrete but about the dynamic process of negotiated sensemaking
Date: 2019-07-14
URL: https://metaphorhacker.net/2019/07/fruit-loops-and-metaphors-metaphors-are-not-about-explaining-the-abstract-through-concrete-but-about-the-dynamic-process-of-negotiated-sensemaking/
Categories: Linguistics, Metaphor, Extended writing
Tags: featured
Note: This is a slightly edited version of a post that first appeared on Medium (https://medium.com/metaphor-hacker/fruit-loops-and-metaphors-metaphors-are-not-about-explaining-the-abstract-through-concrete-but-4e0209dd70b6). It elaborates and exemplifies examples I gave in the more recent posts on metaphor and explanation (/2019/05/explanation-is-an-event-understanding-is-a-process-how-not-to-explain-anything-with-metaphor/) and understanding (/2019/06/5-kinds-of-understanding-and-metaphors-missing-pieces-in-pedagogical-taxonomies/).
One of the less fortunate consequences of the popularity of the conceptual metaphor paradigm (which is also the one I by and large work with on this blog) is the highlighting of the embodied metaphor at the expenses of others. This gives the impression that metaphors are there to explain more abstract concepts in terms of more concrete ones.
Wikipedia: “Conceptual metaphors typically employ a more abstract concept as target and a more concrete or physical concept as their source. For instance, metaphors such as ‘the days [the more abstract or target concept] ahead’ or ‘giving my time’ rely on more concrete concepts, thus expressing time as a path into physical space, or as a substance that can be handled and offered as a gift.“
And it is true that many of the more interesting conceptual metaphors that help us frame the fundamentals of language are projections from a concrete domain to one that we think of as more abstract. We talk about time in terms of space, emotions in terms of heat, thoughts in terms of objects, conversations as physical interactions, etc. We can even deploy this aspect of metaphor in a generative way, for instance when we think of electrons as a crowd of little particles.
But I have come to view this as a very unhelpful perspective on what metaphor is and how it works. Instead, going back to Lakoff’s formulation in Women, Fire, and Dangerous Things, I’d like to propose we think of a metaphor as a principle that helps us give structure to our mental models (or frames). But unlike Lakoff, I like to think of these as an incredibly dynamic and negotiated process rather than as a static part of our mental inventory. And I like to use conceptual intergation or blending as way of thinking about the underlying cogntivive processes.
Metaphor does two things: 1. It helps us (re)structure one conceptual domain by projecting another conceptual domain onto it and 2. In the process of 1, it creates a new conceptual domain that is a blend of the two source domains.
We do not really understand one domain in terms of another through metaphor. We ‘understand’ both domains in different ways. And this helps us create new perspectives which are themselves conceptual domains that can be projected or projected into. (As described by Fauconnier and Turner in The Way We Think).
This makes sense when we look at some of the conventional examples used to illustrate metaphors. “The man is a lion” does not help us understand lesser known or more abstract ‘man’ by using the better known or more concrete ‘lion’. No, we actually know a lot more about men and the specific man we’re thus describing than we do about lions. We are just projecting the domain of ‘lions’ including the conventionalised schemas of bravery and fierceness onto a particular man.
This perspective depends on our conventionalised way of projecting these 2 domains. Comparison between languages illustrates this further. The Czech framing of lions is essentially the same as English but the projection into people also maps lion’s vigour into work to mean ‘hard working’. So you can say “she works as a lion”, meaning she works hard. But in the age of documentaries about lions, a joke subverting the conventionalised mapping also appeared and people sometimes say “I work like a lion. I roar and go take a nap.” This is something that could only emerge as more became conventionally known about lions.
But even more embodied metaphors do not always go in a predictable direction. We often structure affective states in terms of the physical world or bodily states. We talk about ‘being in love’ or ‘love hitting a rocky patch’ or ‘breaking hearts’ (where metonymy also plays a role). But does that really mean that we somehow know less about love than we know about travelling on roads? Love is conventionally seen as less concrete than roads or hearts but here we allow ourselves to be mislead by traditional terminology. The domain of ‘love’ is richly structured and does not ‘feel’ all that abstract to the participants. (I’d prefer to think of ‘love’ as a non-prototypical noun; more prototypical than ‘rationalisation’ but less prototypical than ‘cat’).
Which is why ‘love’ can also be used as the source domain. We can say things like “The camera loves him.” and it is clear what we mean by it. We can talk about physical things “being in harmony” with each other and thus helping us understand them in different ways despite harmony being supposedly more abstract than the things being in harmony.
The conceptual domains that enter into metaphoric relationships are incredibly rich and multifaceted (nothing like the dictionaries or encyclopedias we often model linguistic meaning after). And the most important point of unlikeness is their dynamic nature. They are constantly adapting to the context of the listeners and speakers, never exactly the same from use to use. We have a rich inventory of them at our disposal but by reaching into it, we are also constantly remaking it.
We assume that the words we use have some meanings but it is us who has the meanings. The words and other structures just carry the triggers we use to create meanings in the process of negotiation with the world and our interlocutors.
But this sounds much more mysterious and ineffable than it actually is. These things are completely mundane and they are happening every time we open our mouths or our minds. Here’s a very simple but nevertheless illuminating illustration of the process.
Not too long ago, there were two TV shows that had some premise similarities (Psych and The Mentalist). One of them came out a year earlier and its creators were feeling like their premise was copied by the other one. And they used the following analogy:
“When you go to the cereal aisle in a grocery store, and you see Fruit Loops there. If you look down on the bottom, there’s something that looks just like Fruit Loops, and it’s in a different bag, and it’s called Fruity Loop-Os.” http://th3tvobsessed.blogspot.co.uk/2009/08/psych-vs-mentalist.html (http://th3tvobsessed.blogspot.co.uk/2009/08/psych-vs-mentalist.html)
I was watching both shows at the time but their similarity did not jump out at me. But as soon as I read that comparison it was immediately clear to me what the speaker was trying to say. I could automatically see the projection between the two domains. But even though it seemed the cereal domain was more specific, it actually brought a lot more with it than the specificity of cereal boxes and their placement on store shelves. What it brought over was the abstract relationship between them in quality and value but also many cultural scripts and bits of propositional knowledge associated with cereal brands and their copycats.
But there was even more to it than that. The metaphor does not stop at its first outing (it’s kind of like mushrooms and their rhizomes (https://en.wikipedia.org/wiki/Rhizome) in this way). Whenever, I see a powerful analogy or generative metaphor on the internet, I always look for the comments where people try to reframe it and create new meanings. Something I have been calling ‘frame negotiation (/2012/03/raam-9-abstract-of-doves-and-cocks-collective-negotiation-of-a-metaphoric-seduction/)’. Take almost any salient metaphoric domain projection and you will find that it is only a part in a process of negotiated sense making. This goes beyond people simply disagreeing with each other’s metaphors. It includes the marshalling of complex structuring conceptual phenomena from schemas, rich images, scenarios, scripts, to propositions, definitions, taxonomies and conventionalised collocations.
This blog post and its comments contain almost all of them: http://th3tvobsessed.blogspot.co.uk/2009/08/psych-vs-mentalist.html (http://th3tvobsessed.blogspot.co.uk/2009/08/psych-vs-mentalist.html). First, the post author spends three paragraphs (from third on), comparing the two shows and finding similarities and differences. This may not seem like anything interesting but it reveals that the conceptual blends compressed in the cereal analogy are completely available and can be discussed as if it was a literal statement of fact.
Next, the commenters, who have much less space, return to debating the proposition by recompressing it into more metaphors. These are the first four comments in full:
- Anonymous said… They’re not totally different. It’s more like comparing Fruit Loops to Fruit Squares which happen to taste like beef.
- Nikki0417 said… I think a better comparison would Corn Flakes and Frosted Flakes. Both are made with the same cereal, but one’s sweeter (Psych).
- TV Obsessed said… Sweeter as in more comedy oriented? They are vastly different shows that are different on so many levels.
- Anonymous said… nikki could not be more right with the corn flakes and frosties analogy
Here we see the process of sense making in action. The metaphoric projection is used as one of several structuring devices around which frames are made. Comment 1 opens the the process by bringing in the idea of reframing through other analogs in the cereal domain. 2. continues that process by offering an alternative. 3. challenges the very idea of using these two domains and 4. agrees with 2 as if this were a literal statement but also referring to the metalinguistic tool being used.
The subsequent comments return to comparing the two shows . Some by offering propositions and scenarios, others by marshalling a new analogy.
W. D. Stephenson (https://www.blogger.com/profile/12554378526046963007) said… The reason the Mentalist feels like House is because house is a modern day medical version of Homes as in Holmes Sherlock. Also both Psych and The Mentalist are both Holmsian in creation. That being said I love the wit and humor of psych
Again, there is no evidence of the concrete/abstract duality or even one between less and better known domains. It is all about making sense of the domains in both cognitive and affective ways. Some domains have very shallow projections (partial mappings) such as cornflakes and frosty flakes, others have very deep mappings such as Sherlock Holmes. They are not providing new information or insight in the way we traditionally think of them. Nor are they providing an explanation to the uninitiated. They are giving new structure to the existing knowledge and thus recreating what is known.
The reason I picked such a seemingly mundane example is because all of this is mundane and it’s all part of the same process. One of my disagreements with much of metaphor application is the overlooking of the ‘boring’ bits surrounding the first time a metaphor is used. But metaphors are always a part of a complex textual and discursive patterns (/2018/05/not-ships-in-the-night-metaphor-and-simile-as-process/) and while they are not parasitic on the literal as was the traditional slight against them, they are also not the only thing that goes on when people make sense.
## 5 books on knowledge and expertise: Reading list for exploring the role of knowledge and deliberate practice in the development of expert performance
Date: 2019-06-30
URL: https://metaphorhacker.net/2019/06/5-books-on-knowledge-and-expertise-reading-list-for-exploring-the-role-of-knowledge-and-deliberate-practice-in-the-development-of-expert-performance/
Categories: Education, Framing, Knowledge, Philosophy of Science, Extended writing
Tags: featured
Recently, I've been exploring the notion of explanation (/2019/05/explanation-is-an-event-understanding-is-a-process-how-not-to-explain-anything-with-metaphor/) and understanding (/2019/06/5-kinds-of-understanding-and-metaphors-missing-pieces-in-pedagogical-taxonomies/). I was (partly implicitly) relying on the notion of 'mental representations' as built through deliberate practice. My plan was to write next about how I think we can reconceptualize deliberate practice in such a way that it draws on a richer conception of 'mental representations'. But that is turning out to be a much longer project.
Meanwhile, in a recent conversation about teaching practitioners, somebody mentioned reading Kahneman's 'Thinking Fast and Slow' as being relevant to the problem and we discussed maybe starting a reading group. This got me thinking about what should such a reading group have on its reading list.
The literature on expertise is vast (just look at the Cambridge Handbook of Expertise and Expert Performance (https://www.cambridge.org/gb/academic/subjects/psychology/cognition/cambridge-handbook-expertise-and-expert-performance-2nd-edition?format=HB&isbn=9781316502617)). In my proposed reading list, I would focus on identifying different perspectives on how our mental representations of the world are structured, how we develop them (or how we can help others develop them), how we solve problems with them, and how they are embedded in the social environment in which we function.
## 1. Thinking Fast and Slow by Daniel Kahneman (2011)
Kahneman's famous book (https://en.wikipedia.org/wiki/Thinking,_Fast_and_Slow) is not really focused on experts but rather on the limitations of our thought - summarised under the heuristics and biases banner. But Kahneman's notion of 'System 1' (fast) and 'System 2' (slow) thinking is directly relevant to the question of expertise. Expertise means that one can think about complex issues quickly but also that one can analyze that same issue with deliberate attention to detail. Exactly how this applies to the question of educating experts is a matter of discussion that I think the other books on my list can help elucidate.
## 2. Peak: Secrets from the new science of expertise by Anders Ericsson with Robert Pool (2016)
In this book (https://www.amazon.co.uk/Peak-Secrets-New-Science-Expertise/dp/1847923194), Ericsson (helped by journalist Pool) provides an outline of a cognitive mechanism by which fast thinking is acquired without the sacrifice of deliberation in the concept of 'delibrate practice'. I propose that the key to understanding deliberate practice is not the process of practice but rather on Ericsson's rethinking of the target that the practice should help us achieve. According to Ercisson, what delibrate practice leads is not knowledge or skill but rather 'mental representations'. Mental representations are best thought of as chunks of knowledge (frames, scripts, schemas, etc. - which makes this approach overlap with Kahneman and Tversky's work even though Ericsson does not mention this). This allows experts to perform complex mental operations on very rich subject domains which would be beyond the computational powers of anyone's pure raw intelligence. The best analogy is being able to play chess or speaking a language - this is impossible by simply knowing the rules - we need a rich complex of mental representations to compete at chess or to speak with any fluency.
## 3. The Way We Think: Conceptual Blending and the Mind's Hidden Complexities by Gilles Fauconnier and Mark Turner (2002)
Where Kahneman provides the framework and Ericsson the mechanism of acquisition, Fauconnier and Turner (http://markturner.org/wwt.html) offer us a much more detailed description of the actual structure of 'mental representation' and how it is used during live processing of information. Building on work in cognitive linguistics and semantics, they develop the notion of 'conceptual integration' (or 'blending' as it's more popularly referred to in the field) that explains how multiple 'mental spaces' or 'domains' can be merged seemingly without any conscious effort into new domains (blends) that we can then build further understanding on.
In this context, I'd also recommend reading the parts of Lakoff's 'Women, Fire, and Dangerous Things' that describe what he then called 'Idealized Cognitive Models' and now calls 'frames'. The book is quite vast and not all of it relevant to this question, which is why I wrote a guide to it (/2018/05/how-to-read-women-fire-and-dangerous-things-guide-to-essential-reading-on-human-cognition/).
## 4. Rethinking Expertise by Harry Collins and Robert Evans (2008)
What's missing in all the works I've looked at so far is any awareness of the social embeddedness of expert performance. There is little discussion of types or levels of expertise and barely any mention of how experts interact with one another. In 'Rethinking Expertise (https://www.press.uchicago.edu/ucp/books/book/chicago/R/bo5485769.html)', Collins and Evans propose what they call a 'periodic table of expertise' (which happens to overlap quite nicely with my 5 types of understanding (/2019/06/5-kinds-of-understanding-and-metaphors-missing-pieces-in-pedagogical-taxonomies/)). They think not just about the specialist expert knowledge but also about what they call 'ubiquitous expertise' - all the underlying skills and knowledge required to even get started (such as languages, basic social skills, metacognition, etc.). Most importantly, they also pay attention to 'meta-expertise', i.e. how non-experts evaluate experts and experts judge other experts.
Their notion of expertise relies on the concept of 'tacit knowledge' (later developed by Collins in a separate book (https://www.press.uchicago.edu/ucp/books/book/chicago/T/bo8461024.html)) which is reminiscent of Ericsson's 'mental representations' and echoes Kahneman, as well.
## 5. Reflective Practitioner: How professionals think in action by Donald A. Schön (1983)
While Schön's book (https://www.amazon.co.uk/Reflective-Practitioner-Professionals-Think-Action/dp/0465068782)has had a profound impact in terms of citation and ways of thinking, I suggest that it has been largely under-appreciated for its depth of epistemological insight. Despite being more than 2 decades older than any of the other books on this list, it is very much still relevant. It considers the very nature of 'practical knowledge' as opposed to 'academic knowledge'. Schön, more than any of the others thinks about the practical needs of a person needing to achieve practical tasks with their knowledge in a complex situation. He highlights the tensions between the technical preparation of experts that focuses on knowledge about a subject and the practical needs of a practitioner who needs to act in such a way that simply recalling information would not be sufficient. His concept of 'reflection-in-action' could be seen as a precursor or better still a companion to the notion of 'deliberate practice'.
Schön followed this up with Educating The Reflective Practitioner (https://www.amazon.co.uk/Educating-Reflective-Practitioner-Professions-Education-ebook/dp/B0022NGE62) which focuses on the practical question of structuring a training course. Another reason to include Schön on this list is that he focuses more directly on 'professional' expertise.
## Bringing it all together
What these books have in common is an underlying conception of knowledge and its processing. But what they lack is almost any awareness of each other. This makes them add up to more than just the sum of their parts.
Kahneman mentions Ericsson in a footnote and Ericsson and Collins appear jointly in the Cambridge Handbook I mentioned at the start. But they largely travel in separate spheres. Bizarrely, none of them refers to Schön. And all of them are completely unaware of Fauconnier and Turner, who in turn ignore the work done outside their field of cognition (even though we can trace the lineage of their work on cognitive domains directly to Schön's earlier work on metaphor).
All these approaches are clearly converging on the same thing but they don't do it using the same terminology, methods or even a shared conceptual framework. Which is why reading just any one of them would probably not be enough to get at the full scope of the issues involved.
I'm not certain that this selection is the most representative of the field. It is certainly not exhaustive and it is definitely shaped by my idiosyncratic intellectual journey and personal interests. But my hope is that it does triangulate the problem domain in a way that a more narrowly focused selection would not.
## Writing as translation and translation as commitment: Why is (academic) writing so hard?
Date: 2019-06-15
URL: https://metaphorhacker.net/2019/06/writing-as-translation-and-translation-as-commitment-why-is-academic-writing-so-hard/
Categories: Education, Linguistics, Metaphor, Extended writing, Writing
Tags: featured
##
This book will perhaps only be understood by those who have themselves already thought the thoughts which are expressed in it—or similar thoughts. It is therefore not a text-book. Its object would be attained if there were one person who read it with understanding and to whom it afforded pleasure.
(opening sentence of the preface to Tractatus Logico Philosophicus by Lugwig Wittgenstein, 1918)
## Background
I’ve recently been commenting quite a lot on the excellent academic writing blog (which I mostly read for the epistemology) Inframethodology (https://blog.cbs.dk/inframethodology) by Thomas Basbøll. Thomas and I disagree on a lot of details but we have a very similar approach to formulating questions about knowledge and its expression.
The recent discussion was around the problem of ‘writing as expressing what you know’. While I find it very useful to distinguish between writing to describe what you know and writing to explore and discover new ideas (something I first reflected on after reading Inframethodology), I commented:
I still find that no matter how well I think I know my subject, I discover new things by trying to write it down (at least with anything worth writing).
Thomas responded in a separate blogpost (https://blog.cbs.dk/inframethodology/?p=2518), first picking up on my parenthetical:
Can it really be true that the straightforward representation of a known fact is not “worth writing”? Is the value of writing always to be discovered (by way of discovering something new in the moment of writing)? I think Dominik is thinking of kinds of writing that are indeed very valuable because they present ideas that move our own thinking forward and, ideally, contribute positively to the thinking of our peers. But I also think there is value is writing that doesn’t do this, writing that is, for lack of a better word, boring.
With this, I agree wholeheartedly. 110% coach! Yes, this was a throwaway line I wasn’t comfortable with even as I was writing it. The majority of my writing is mundane: emails, instruction manuals, project proposals, etc. They may or may not be “worthy” but they certainly have a worth. And people who do nothing but that sort of writing certainly do not do anything I would find ‘beneath me’ or not worthy. I might have been better served by the term ‘quotidian’ or even ‘instrumental’ writing.
I agree even more with Thomas’s elaboration (my emphasis):
In fact, I think it’s the primary of value of academic writing and one of the reasons that so many people (and even academics themselves) almost equate “academic” (adj.) with “boring”.The business of scholarship is not to bring new ideas into the world, indeed, the function of distinctively academic work (in contrast to, say, scientific or philosophical or literary work) is not to innovate or discover but to critique, to expose ideas to criticism. In order for this happen efficiently and regularly, academics must spend some of their time representing ideas that are not especially exciting to them along with their grounds for entertaining them. They must present their beliefs to their peers along with their justification for thinking they’re true. And they must do this honestly, which is to say, they must not invent new beliefs or new reasons for holding them in the moment of writing. They must write down, not what they’re thinking right now, but what they’ve been thinking all along.
I find this an incredibly valuable perspective and when I think of my own writing, I think this is precisely where I’ve often been going wrong. This is partly because academic writing is more of a hobby than a job, so I don’t have the time to do more than write to discover. But it is partly because of my temperament. I don’t enjoy the boring duties of writing things I know down and then formatting them for the submission to a journal. I prefer to work with editors which is why the bulk of my published writing is in journalism or book chapters.
But there is still another aspect that needs to be explored. And that is, why do most people find it so difficult to write down what they know even while taking into account all of the above.
## Writing as translation
I propose that a good way to think about the difficulty of writing to describe our thoughts is to use the metaphor of translation. We can then think of the content of our thoughts in our head as a series of propositions expressed in some kind of ‘mentalese’. And when we come to write them down, we are essentially translating them into ‘writtenese’ or in this case, one of its dialects ‘academic writtenese’.
This is made more complicated by the existence of a third language - let’s call it ‘spokenese’. We are all natively bilingual in ‘mentalese’ and ‘spokenese’ even if not everybody is very good at translating between these two languages. In fact, children find it very difficult until quite late ages (10 and up) to coherently express what they think and even many adults never achieve great facility with this. Just like many natively bilingual speakers are not very good at translating between their two languages.
But nobody is a native speaker of ‘writtenese’. Everybody had to learn it in school with all its weird conventions and specific processing requirements. It is not too outlandish to say (and I owe this to the linguist Jim Miller) that writing is like a foreign language. (Note: see some important qualifications below).
When we are translating from mentalese to academic writtenese, we are facing many of the same problems translators of very different languages faces. The one I want to focus on is ‘making commitments’.
## Translation as commitment: Making the implicit explicit
Perhaps the most difficult problem for a translator (I speak as someone who has translated hundreds of thousands of words) is the issue of being forced by the way the target language operates to commit to meanings in the translation where the structure of the source language left more options for interpretation.
Let’s take a simple paragraph consisting of three sentences (Note: this is a paraphrase of an example given by Czech-Finnish translator at a conference I attended some years ago):
The prime minister committed to pursue a dialogue with the opposition. This was after the opposition leader complained about not being involved. She confirmed that he would have a seat at the table in the upcoming negotiations.
The first commitments I have to make at some point is to the gender of the participants in the actions I write about. In English, I can leave the gender ambiguous until the third sentence. In Finnish, which does not have gendered third-person-singular pronouns, I don’t have to express the gender at all.
In Czech (and many other languages), on the other hand, I have to know the gender of the prime minister from the very first word. Like actor and actress in English, all nouns describing professions have built-in genders (this is not optional as in English because all Czech nouns have assigned some grammatical gender). I also need to express gender as part of the past tense morphology of all verbs. So even if I could skirt the gender of the ‘leader’ (there are some gender-ambiguous nouns in Czech), I would have to immediately commit to it with the verb ‘complained’. Which is why knowledge of their subject is essential to simultaneous translators.
But this is a relatively simple problem that can be solved by reference to known facts about the world. A much more significant issue is the differential completion of certain schemas associated with types of expressions. Let’s take the phrase ‘committed to pursue’. The closest translation to the word ‘commit’ is ‘zavázat se’ which unfortunately has the root ‘bind’. It is therefore ever so slightly more ‘binding’ than ‘commit’. I can also look into something like ‘promise’ which of course is precisely what the prime minister did not do.
Then, there is the word ‘pursue’. One way to translate it is ‘usilovat o’ which has connotations of ‘struggle to’. So ‘usilovat o dialog’ is in the neighborhood of ‘pursue a dialog’ but lacks the sense of forward motion making it seem slightly less like the dialog is going to happen. So here each language is making subtly different commitments.
When you’re translating academic writing, there are hundreds of similar examples, where you have to fill in blanks and make some claims seem stronger and others weaker. And even if you know the subject intimately (which I did in most cases), you often have to insert your judgement and interpretation. And the more you do that, the less certain you feel that you got the meaning of the original exactly right. This is even when while reading the original, I had no sense of something being left unexpressed. The only way to get this right is to ask the author. But even that may not always work because they may not remember their exact mental disposition at the time of writing.
## Writing as filling in holes in our mind
I believe that this is exactly the experience we have when we write about something that only exists in our head or something we’ve only previously talked about. Even when I’ve given talks at conferences and had many conversations with colleagues, writing my ideas down remains a difficult task.
When writing, the structure of ‘writtenese’ (as well as the demands of its particular medium) forces me to make certain commitments I never had to make in ‘mentalese’ (or even ‘spokenese’). I have to fill out schemas with detail that never seemed necessary. I have to make more commitments to the linearity of arguments, that could previously run parallel in my head. So when I write it is not clear what should come first and what last.
When I just write down what’s in my head (or as close to it as it is possible), it is unlikely to make any sense to anybody. Often including myself after some time. I need to translate it in such a way that all the necessary background is filled out. I also need to use the instruments of cohesion to restore coherence to the written text that I felt in my mind without any formal mental structure.
But during this process, I often become less certain. The act of writing things down triggers other associations and all of a sudden I literally see things from a different perspective. And this is often not a comfortable experience. Many writers find this a source of great stress.
This is, of course, true even of writing instructions and directions. Often, when describing a process, we find there are gaps in it. And when writing down directions, we come to realise that we may not know all aspects of the familiar sufficiently well to mediate the experience to someone else.
## Teaching writing as translation
Translation is a skill that requires a lot of training and practice. In many ways, a translator needs to know more about both languages than a native speaker of either. And then they need to know about different ways of finding equivalent expressions between the two languages in such a way that the content expressed in the source language produces similar mental effects when reading in the target language. This is not easy. In fact, it is frequently impossible to achieve perfectly.
When I translate I often refer to a dictionary (such as slovnik.cz) that lists as many possible alternatives of words even if I know exactly what the original ‘means’. This is because I want to see multiple options of expressing something which may not be immediately triggered by my understanding of the whole.
But for this to work, I need to have done a lot of deliberate reading in both languages to know how they tend to express similar things. At the early stages, I may approach this more simply as learning to speak a language. I may learn that ‘commit to pursue’ is best translated as ‘zavázat se usilovat o’. But I have to back that up by a lot of reading in both languages, studying other translators’ work and making hypotheses about both languages and the differences between them. Eventually, this becomes second nature and to translate fluently, we need to ‘forget’ the rules and ‘just do it’.
So how could we apply this to teaching (academic) writing? We need to start by ensuring that students have enough facility in both the source and the target languages. We usually assume greater fluency in the source language (most translators work primarily in the direction of native to non-native). So in this case, we need to focus on the structures and ways of ‘academic writtenese’.
We can very much approach this as teaching a foreign language. Our first aim should be to help students acquire fluency in the language of academic writing. We need to give them some target structures to learn. This should ideally be based on an actual analysis of that writing rather than focusing on random salient features. But ultimately, the key element here is practice.
Then we also need to focus on helping the students develop better awareness of their native mentalese and how to best map its structures onto the structures of writtenese. We can do this by helping them write outlines, create mind maps, come up with relevant key words, and of course, read a lot of other people’s writing, think about it, and then write summaries in similar ways.
None of these are particularly revolutionary ideas and they are being used by writing teachers all over the world. What I’m hoping to do here is to provide a metaphor to help focus the efforts on particular aspects of what makes the translation from thought to writing difficult.
## Writing as playing a musical instrument
One final analogy that can help us here is the idea of writing as playing a musical instrument. This analogy is in many ways even more apt. When we play a musical instrument, we are initially translating relatively vague musical ideas into actual notes (melodies and harmonies) by way of the structures given to us by the musical instrument.
We may start by learning some chords to accompany a song we hear but later we will progress into more details of musical theory which will allow us to express more elaborate ideas. But, in fact, this also allows us to have more those more elaborate ideas in the first place.
Initially, our ability to express musical ideas via an instrument (such as piano or guitar) will be limited by our skill. We may not even realize what exactly the idea in our head was until we’ve played it. And often, what we can play limits the ideas we have. Jazz teachers often say something like ‘sing your solos first and then play’ (others call it ‘audiation’). But this is not trivial and requires extensive training. Which is why one common advice for jazz musicians is to transcribe (or at least copy) famous songs and solos. But as you’re transcribing and copying, you’re supposed to notice patterns in how musical ideas are expressed. You can then recombine them to express what is in your ‘musical mind’.
But it seems that the musical ideas and their form of expression are never completely separate. They are not a pure translation but rather a co-creation. And this is true of any good translation and probably also ultimately true about any act of writing. We are using a different medium to express an existing idea but in the process, we are filling gaps in the ideas, creating new connections until we ultimately cannot be completely certain which came first.
As we get better at translation, music or writing, there are some levels about which the last part does not hold true. There are some ideas we can truly and faithfully translate from our head to paper, musical instrument or from one language to another. This is why practice is so important. But at the highest levels of difficulty, writing, translation and music making will always be acts of co-creation between the medium and the message.
## Teaching writing as music
So finally, could we teach writing in the same way as we teach music? We certainly could. Just like teaching a foreign language, teaching music is mostly dependent on a lot of practice.
But perhaps there are some techniques that music teachers use that could be useful for both language teachers, translators and writing coaches.
One is the emphasis on patterns. The idea of practicing scales, licks, or chords relentlessly (up to hours a day) holds a lot of appeal. Perhaps we start teaching self-expression with writing too soon. Maybe we should give students some practice patterns to repeat in different combinations. Then we could tell them to just copy and then dissect parts of good texts. The idea of ‘mindless’ copying will probably stick in many teachers’ craws. But just analysing reading will never be enough. Students need the experience of writing some good writing. If only to develop some muscle memory. And while it should never be completely mindless, it should also perhaps not be completely meaningful from the very start. Of course, we could invent numerous variations on this approach to transform the texts in various fun ways while still making sure, students are writing extended chunks and developing fluency. The point is that we would not be focusing on self-expression but developing a language for self-expression.
Music teachers and students use what has been described by Anders Ericsson (https://www.amazon.co.uk/Peak-Secrets-New-Science-Expertise/dp/0544456238) as ‘deliberate practice’. Ericsson gives the example of Benjamin Franklin who used similar techniques to improve his writing:
He first set out to see how closely he could reproduce the sentences in an article once he had forgotten their exact wording. So he chose several of the articles whose writing he admired and wrote down short descriptions of the content of each sentence—just enough to remind him what the sentence was about. After several days he tried to reproduce the articles from the hints he had written down. His goal was not so much to produce a word-for-word replica of the articles as to create his own articles that were as detailed and well written as the original. Having written his reproductions, he went back to the original articles, compared them with his own efforts, and corrected his versions where necessary. This taught him to express ideas clearly and cogently.
Obviously, this was not all there was to it, but it is very much reminiscent of what music students do. It seems to me that most beginner writers are often asked to do too much at the very start and they never get a chance to improve because they essentially give up too soon.
## Writing is NOT foreign language, translation or music: The Unmetaphor
Writing is writing! It has its specific properties that we need to attend to if we want to see all of its complexities. We must use metaphors to help us do this but always by remembering that metaphors hide as much as they reveal. One useful way of understanding something is to create a sort of unmetaphor: a listing of similar things that are different from it in various respects. This is something that, while not uncommon, is done much less than it should be when using analogies.
### Written language is not a foreign language
Some of the fundamental mental orientations of a language are shared between the written and spoken forms. This includes tense, aspect, modality, definiteness, case morphology, word categories, meanings of most function words, the shape of words, etc. These present some of the most significant difficulties to learners of foreign languages making it very difficult to acquire a second language by exposure alone after a certain age for most adults.
Writing, on the other hand, can be acquired predominantly by exposure alone for many (if not most) adults. There are many people who acquire native-like competence in the written code in the same way they acquired their spoken language competence (even if there are just as many who never do). And we must also be mindful (as Douglas Biber's research revealed) that there is a bigger difference between some written genres then there is between writing and speech overall. So we should perhaps attend to that.
### Writing is not translation
That writing is not actually translation is contained in the fact that written language is not actually a foreign language. There are many genres and registers in any language with their specific codes. And we could call going from one code to another translation much more easily than going from what I called ‘mentalese’ and ‘writtenese’. (Again, the work of Douglas Biber should be the first port of call for anyone interested in this aspect of writing.)
But most importantly, what I called ‘mentalese’ does not actually have the form of a language. Individuals differ in how they represent thoughts that end up being represented by very similar sentences. Some people rely on images, others on words. For some, the mental images more schematic and for others, they have more filled in details. For instance, Lakoff asked how different people imagine the ‘hand’ in ‘Keep somebody’s at arm’s length’. And the responses he got were that for some the hand is oriented with the palm out, others with the palm in. For some, it includes a sleeve, for others it does not. Etc.
### Writing is not music
I’ve already written about the 8 ways in which language is not like music (/2018/03/10-ways-in-which-music-is-like-language-and-8-more-important-ways-in-which-it-is-not/). And they all apply to writing, as well. The key difference for us here is that music cannot express propositions. This means that musical expression can be a lot freer than expressing ideas through writing.
We could argue that writing is more like music than spoken language because it requires some kind of an instrument. Pen, paper, computer, etc. But we usually learn these independently of the skill of expressing ourselves through writing. My ability to play the piano is much more closely tied to my ability to express my musical meanings. However, people write just as expressive prose by the hunt and peck method as when they touch type. One can even dictate a ‘written text’ - that’s how independent it is of the method of production.
Of course, improving one’s facility with the tools of production can improve the writing output just by removing barriers. This is why students are well-advised to learn to touch type or to use a speech-to-text method if they struggle for other reasons (e.g. visual impairment or dyslexia). But when it comes down to it, this is just writing down words and as we established, writing in most senses is more than that.
## Conclusions and limitations
Ultimately, writing and translation are not the same. Just as writing and music are not the same. But there are enough similarities to make it worthwhile learning from each other.
Many writers have developed great skills by the ‘tried and tested’ approach of ‘just doing it’. But we also know that even many people who do write a lot never become very ‘good’ at it. They struggle with the mechanics, ability to express cogently what’s in their minds, or just hate everything about it.
For some beginner writers, the worst thing we could do is give them a lot of mindless exercises. These people will want to do it first and would hate to be held back. Just like many students of languages or music like dive off the deep end. But equally, for many others, telling them to ‘just do it’ is the perfect recipe for developing an inferiority complex or downright phobias of writing.
But all of these writers will need lots of practice - regardless of whether we provide lots of ladders and scaffolding or just put a trampoline next to the edifice of their skill. In this, writing is exactly like music, language and translation. You can only get better at it by doing it. A lot!
I started with a quote from Wittgenstein. But he also famously said in summarising his book:
What can be said at all can be said clearly; and whereof one cannot speak thereof one must be silent.
I think we saw here that this is not necessarily how the act of writing presents itself to most people.
He then continued:
The book will, therefore, draw a limit to thinking, or rather—not to thinking, but to the expression of thoughts; for, in order to draw a limit to thinking we should have to be able to think both sides of this limit (we should therefore have to be able to think what cannot be thought). The limit can, therefore, only be drawn in language and what lies on the other side of the limit will be simply nonsense.
This is was the so-called “early Wittgenstein” before the language games and family resemblances. He spent the rest of his career unpicking this boundary of sense and non-sense. Coming to terms with the fact that what is thought and what is its expression are not straightforward matters.
So all the metaphors notwithstanding, we should be mindful of the constant tensions involved in the writing process and be compassionate with those who struggle to navigate them.
## 5 kinds of understanding and metaphors: Missing pieces in pedagogical taxonomies
Date: 2019-06-14
URL: https://metaphorhacker.net/2019/06/5-kinds-of-understanding-and-metaphors-missing-pieces-in-pedagogical-taxonomies/
Categories: Education, Knowledge, Metaphor, Extended writing
Tags: featured
## TL;DR
This post outlines 5 levels or types of understanding to help us better to think about the role of metaphor in explanation (/2019/05/explanation-is-an-event-understanding-is-a-process-how-not-to-explain-anything-with-metaphor/):
- Associative understanding: Place a concept in context without any understanding.
- Dictionary understanding: Repeat definitions, give examples, and make basic connections.
- Inferential understanding: Make useful inferences based on knowledge about - but without ability to use the understanding in practice. Requires more than just one concept.
- Instrumental understanding: Use the understanding as part of work in a field of expertise. Impossible to acquire for an isolated concept.
- Creative understanding: Transform understanding of one domain by importing elements from another. Requires instrumental understanding - goes beyond hints and hunches.
## Introduction
In a previous post (/2019/05/explanation-is-an-event-understanding-is-a-process-how-not-to-explain-anything-with-metaphor/), I proposed three uses of metaphor leading to different levels of understanding.
- Metaphor as invitation
- Metaphor as a tool
- Metaphor as catalyst
Only 2 and 3 led to any meaningful understanding and that could only be achieved by acquiring some ‘native’ structure of the target domain. But I was rather loose with how I used the word ‘understanding’. I was using notions like ‘meaningful understanding’ or ‘useful understanding’ but never went into any detail. That is the purpose of this post.
In what follows, I provide a sketch for one way of classifying different kinds of understanding. They are not meant to be descriptions or even discovery of some sort of 'natural kinds'. Instead, I find them to be a useful way of looking at understanding from the perspective of metaphoric cognition.
## Associative understanding
Associative understanding is the ability to place something in a context or category without necessarily knowing almost anything about it. So, we may know that an emu is a flightless bird without knowing anything else about it. We could also think of this kind of understanding as a vague notion.
This is the kind of understanding the vast majority of education leaves us with after a few years. Watching a documentary, a TV quiz show, or reading a popular news article fosters this kind of understanding.
Many people can get very far with displaying this kind of understanding - such as con artists impersonating doctors - by successfully imitating experts. The famous Sokal hoax was based on the same principle: making plausible sounding noises can get you published in a prestigious publication. But it is even possible to pass a poorly constructed multiple choice knowledge test with just this understanding by being able to eliminate the wrong options rather than by knowing the correct ones.
The associations can be of various kinds. They can be in the form of basic-category labels (such as - this is an animal). They could place the thing into a discipline - such as ‘something they do in chemistry’. And they could simply be in the form of ‘this is the thing that my friend always talks about’. Or they could also just be a part of the cultural vocabulary without a proper object of understanding.
For example, in the 1960s' Czechoslovakia there was a famous pop song called ‘Pták Rosomák’ (The Bird Wolverine). The band simply liked the sound of the Czech word for ‘wolverine’ and its rhyme with the word for ‘bird’. Wolverines are not native to Europe or well known outside of this song. I did not find out what the word meant until I learned it in English (I also knew what the English word wolverine meant long before I looked it up in a Czech dictionary). When I presented this at a conference on cognition in Prague, most Czech academics in the audience were surprised by the meaning. Yet, if you asked them - do you understand the word ‘rosomák’, they would have said ‘of course, I do’. But it was just an associative understanding.
My claim is that the vast majority of what passes for understanding and knowledge in ‘polite society’ is of the associative kind. People feel comfortable when concepts like evolution or philosophy are mentioned but have only the vaguest idea of where they belong.
My favourite example of this is Monty Python’s ‘Philosopher’s song (https://www.youtube.com/watch?v=PtgKkifJ0Pw)‘. All the audience needs to know to appreciate the jokes is that there is a philosopher stereotype and that certain names are of philosophers. In fact, by their own admission (citation needed but I did hear it in an interview), the authors of these sketches also did not know much more than the names. Even the little nod to knowledge in ‘John Stuart Mill of his own free will’ is just a glimmer of something deeper.
Associative understanding is pretty much only useful for social signalling. It can also play a role in making a new field appear more familiar in later stages. I have had that experience several times when vague memories from school made me feel more confident I was on the right track when I set about studying a subject in depth even if I had very little more than a vague feeling about something. But on its own, this kind of understanding has little practical value.
In formal instruction, we generally start with the next step but over time, without practice, this is the kind of understanding, we’re left with. But in literature on pedagogy, it is mostly unaddressed. It is the kind of understanding below the bottom rung of Bloom’s taxonomy. But many teachers encounter it when at the end of classes students come and ask questions that barely show a hint of an understanding that makes it seem like they may not have even been in the same room.
## Lexical understanding
At this level, we can repeat a definition as we might find it in a dictionary and give a few examples. We can look at a picture and say, this is an emu. It lives in Australia and it is a kind of ostrich. For something like an emu, it may well be enough for most of us.
This is the kind of understanding we may be able to take away from a quick explanation of something. It is the sort of understanding most tests check for. It is also often used as a proxy for intelligence or ‘being smart’. Lexical understanding is what is required of successful quiz show panellists. UK shows such as ‘Mastermind’, ‘Brain of Britain’ or ‘University Challenge’ are great examples of these.
Conversely, lack of lexical (and sometimes even associative) understanding is also often given as an example of educational decline or lack of intelligence.
This would be roughly equivalent to the ‘Knowledge’ and ‘Comprehension’ levels on the Bloom’s taxonomy. It is the minimum target for instruction but it is very unstable. Unless it has been recently used, it often reverts to associative kind of understanding.
This kind of understanding is generally not very useful outside the educational context. This is the kind of understanding that is the result of ‘teaching to the test’. It can be leveraged into something more but only with practice and application.
In terms, of frames or mental representations, we could say that the only mental representations developed as part of this understanding are propositional or rich imagery. Meaning, we have sentences or images in our head that we can draw on but we would find it very hard to combine them into larger wholes.
This level and the transition from this level to the next are where what we call pedagogy plays the most important role.
## Inferential understanding
This kind of understanding lets us make useful inferences about the concept in context. It requires some knowledge of a whole domain or several domains. You can never understand a solitary concept at this level. But it does not necessarily require deep ability or skill. I know nothing about emus, so I cannot think of an example that would not be misleadingly trivial.
But I have a personal example from when I was recently catching up on the latest developments in machine learning. I was reading about different types of neural nets. And when I was reading about CNNs (Convolutional Neural Networks) which are usually used for images, I had an idea for using the similar approach to process language by representing text in a way similar to the way images are represented. And it turned out there are already papers and models out there that do just that.
Inferential understanding is the kind of understanding that good students develop about favorite subjects that they pursue later. The kind of understanding that collaborators develop about each others’ discipline in interdisciplinary projects. The kind of understanding good generalist managers develop about the domains in which they supervise subject experts. Or really good journalists develop about areas on which they report. This is also the kind of understanding experts have about related fields or that teachers have about some of the more advanced areas of their field.
The sociologist of science Harry Collins described in one of his books (I think it was ‘Rethinking Expertise’) how he could pass some knowledge tests in gravitational wave physics better than professional physicists from adjacent specialisations. This was after many years of observing these physicists but without any real ability to the actual calculations or research required.
It may not always be easy to tell the boundary between this and lexical or even associative understanding. This is the kind of understanding potentially displayed by an audience member at a lecture who asks a question that is then described as ‘a good question’ by the presenter. But often this is just a fluke. A random hit based on superficial resemblance of words in a definition.
This is the kind of understanding that sort of ‘does not count’ in the terms of Bloom’s hierarchy. We feel it is insufficient because it is not something people consciously aim at in instruction. But it is in many ways the best we can hope for. It is the first kind of any useful knowledge.
It requires more developed mental representations. Representations where the propositions and rich images are replaced by schemas and scenarios. These are a sort of useful compressions that can be blended (or integrated) with others. What it means that when reasoning with these concepts, we can use them as whole units (mental chunks) rather than laboriously compute them from first principles.
It may also derive from some basic level of instrumental understanding. The humour in XKCD cartoons can be understood with a combination of inferential and instrumental understanding. I immediately understood this comic famous among programmers (https://xkcd.com/327) without being a programmer myself but having some skills with databases and knowledge of common problems with security.
But for the most part, we cannot use this understanding for actual work. This is where the humanities and sciences often diverge. It is possible to pretend (even to oneself) that this understanding lets us do real useful work in history or sociology. Whereas with mathematics, engineering, medicine, or biology, the barrier between this and instrumental understanding is much more clearly defined by specialised tools such as mathematics and chemistry. But if we look at the many former physicists or biologists who have tried their hand at philosophy, sociology or even literary criticism, we see that even here, this kind of understanding is not enough.
You really need more to have a chance of doing something useful.
## Instrumental understanding
This is the kind of understanding experts and practitioners have. It requires being able to use the concepts or tools in practice. I don’t have any instrumental understanding of convolutional neural networks. I couldn’t build one and possibly couldn’t even reconstruct the exact way in which it works.
This level of understanding or ability or skills requires more than just reading or learning about. It requires practice and building of mental representations which only comes from long-term engagement with a subject. For example, I don’t have that kind of understanding of neural nets, but I do have it of metaphor.
I can create metaphors, identify them in text, speak to the controversies around them, compare and contrast the various theories of metaphor. I can teach somebody how metaphors work. I can write a successful paper or give a conference presentation in the field. If somebody wants to know about metaphor they can come to me. Other people with good instrumental understanding of metaphor may disagree with some of what I have to say, but they won’t do it (I hope) as they would with somebody who has just an associative, lexical or even inferential level of understanding - e.g. knowing that metaphor has something to do with poetry. You have to put in the work.
This work may require actual repetitive practice (such as working out math problems or analysing texts). It absolutely requires extensive engagement with other experts in the field. Taking classes, going to conferences, reading latest research, writing papers, blogs, etc. That’s why loner autodidacts almost never reach this level of understanding.
Here the distinction between understanding and ability or skill becomes blurred. Mental representations develop at the highest levels of schematicity. This means that an expert can look at a very complex situation and treat it as one unit that can be blended with other complex units in a way that only the relevant parts are engaged.
For instance, I can read a complex argument about metaphor and immediately compare it with three other complex arguments about metaphor - not because I have a large mental capacity for abstract concepts but because I have developed a number of highly schematic mental representations about the shapes of arguments people make about metaphor. This way, I can project these schemas onto the argument as one big chunk.
Perhaps an even better analogy is learning a foreign language. I may know all the rules and words but I cannot speak the language with any level of fluency until I have developed larger chunks I can just slightly modify. It is simply impossible for even the most highly mentally endowed human to dredge up individual words, apply rules to them and combine them into a sentence quickly enough to speak with any level of coherence. It’s even worse for understanding. Just reading a text with a dictionary is such a slow affair that we forget what a sentence was about before we get to the end.
In other words, we can then define instrumental understanding as developing a basic fluency in the language of the discipline. And this takes time, targetted practice, and active ‘communicative’ engagement across a whole field.
In the ‘hard sciences,’ it requires a good facility with formalisms or even equipment and in the ‘softer’ disciplines it relies on extensive reading, talking, and writing.
Here we are at a much wider aperture of our knowledge funnel. It is therefore impossible to exactly compare 2 people’s levels of instrumental understanding. Everybody will have a slightly different set of mental representations. Also, many people will only be able to ‘perform’ at this level some of the time or only for small chunks of their discipline.
At this level, pedagogy is much less relevant. This is where it makes a lot less sense to talk about teaching and learning if only because it is impossible to acquire this level of understanding purely in the classroom. Training, coaching or even apprenticeship are much better models.
## Creative understanding
Creative understanding is instrumental understanding with a transformative element. This requires knowledge of several domains and their creative intermingling. It is the sort of understanding innovators in their field have. This can lead us to a complete rejection of the thing we understand as an independent concept.
For example, I have long argued that metaphor is not the only place in language where domain projection occurs and that we should not think of it as something special but rather as a shortcut for thinking about broader phenomena of framing or cognitive models. I found this a useful way of extending the concept. So, I can make a serious statement such as ‘metaphor and metonymy are the same thing’ that can be productive in the study of metaphor. But it only makes sense because I can actually distinguish between metaphors, similies (/2018/05/not-ships-in-the-night-metaphor-and-simile-as-process/), synechdoches or metonymies (/2013/12/binders-full-of-women-with-mighty-pens-what-is-metonymy/) in an instrumental way, and I can also reproduce arguments that maintain that the difference between metaphor and metonymy is crucial for understanding figurative language.
It is hard to say whether this type of understanding is even a part of the funnel hierarchy. Perhaps it is just an ingredient (catalyst) to instrumental understanding. But I do want to stress that it only works as a catalyst to instrumental understanding. As I showed in my post on types of metaphors, creativity needs to start from somewhere.
We may often confuse almost accidental insights by people with inferential or even just lexical understanding for creativity. But this is like recognising a melody in the sounds a child makes by randomly banging on the piano keyboard.
We often valorise the outsider perspective in a field. And it certainly can act as a catalyst for creativity but only if it has proper instrumental understanding to lean on.
## Conclusions and limitations
I cannot stress enough that this classification is just a useful heuristic. I am not claiming that this kind of classification of understanding is exhaustive or even that it represents some sort of a natural category. But I found it useful when thinking about explanations and pedagogy.
## Approaches to classifying understanding
It is quite common to distinguish between shallow and deep understanding. This is intuitively obvious but not very helpful because it assumes the existence of some sort of objective scale of a depth of understanding.
We can also distinguish understanding from knowledge for example by differentiating between explicit and tacit knowledge. Understanding and explicit knowledge intuitively overlap even if we don't have a firm definition of either. If we understand something, we can mentally manipulate it and, most importantly, pass it along.
But the boundaries between tacit and explicit knowledge are not firm. All explicit knowledge depends on some tacit knowledge - or in other words, all understanding depends on knowledge. We could even say that deep learning is the process of transforming understanding into knowledge. In the sense, that we need to build up schematic mental representations to be able to manipulate ever more complex combinations of concepts.
Another way to try to get at understanding is to investigate how to achieve it. Bloom’s taxonomy of educational objectives is one famous example. There are many tweaks and elaborations - some as extreme as Jack Koumi’s 33 pedagogic roles. But they are ultimately not very satisfying because they already assume we know what the understanding is.
## Understandings as a process revisited: The wave and the funnel
Even though these different types of understanding are ‘broadly hierarchical’, I want the emphasis to be on ‘broadly’. It would make no sense to think of these as a straightforward linear hierarchy measurable on a scale of discrete and comparable units. They are more like overlapping waves. Layers of water covering the beach in successive bursts as the tide is coming in.
But that metaphor does not make it easy to visualise the differences and mutual interdependence. It only evokes how hard and unreliable it is to do so. But for the purposes of this comparison, I’d like to offer something more like a funnel (which I also brought up in the context of the metaphor explanation hierarchy (/2019/05/explanation-is-an-event-understanding-is-a-process-how-not-to-explain-anything-with-metaphor/)) or inverted cone.
The substance that fills the funnel might be a mixture of effort and coverage of material. This makes it easy to visualise the fact that it takes much more effort, time and background knowledge to get from level 3 to level 4 than it does to get from level 1 to level 2. Also, at the higher levels, the concepts themselves transform and interconnect. So it is not possible to understand them in isolation.
This truly takes into account the processual nature of understanding. The funnel also needs to be constantly topped up to maintain certain levels. But it can also underscore the fact that we can never perfectly compare 2 people’s levels of understanding. Because at the higher levels, the funnel is so broad, not everybody will have filled it in the same way with exactly the same substance.
I got this idea from ACTFL language competency levels and I think it is one of the most underappreciated metaphors in education.
Another really useful thing ACTFL does is that it defines low, mid and high sublevels for each competency levels. And a part of the definition of the ‘high’ sublevel is that the person can function at the ‘low’ sublevel of the next level about half the time. (E.g. a Novice-Low can function as Intermediate-Low about 50% of the time). During the test (most often an interview), the examiner establishes a floor and a ceiling rather than pinpointing an exact point on a scale.
This very much applies to the levels in my metaphor. There are no clear boundaries between these levels of understandings. In as much as they are levels in the first place.
## Explanation is an event, understanding is a process: How (not) to explain anything with metaphor
Date: 2019-05-29
URL: https://metaphorhacker.net/2019/05/explanation-is-an-event-understanding-is-a-process-how-not-to-explain-anything-with-metaphor/
Categories: Education, Framing, Knowledge, Metaphor, Extended writing
Tags: featured
## TL;DR
- There are at least 3 uses of metaphor in the educational process: 1. Invitation to enter; 2. An instrument to grasp knowledge with; 3. Catalyst to transform understanding. Many educators assume that 1 is enough but it rarely leads to any useful understanding.
- Explanation is a salient part of the educational process to such an extent that it is often allowed to stand for all of it even though it is only one step.
- Explanation often helps the person doing the explaining more than the person being explained at.
- Metaphors and explanations have been misused by educators from Socrates to Rousseau.
- A metaphor can only be successful if the student already has some knowledge of the target domain. Knowledge of the source domain is often less important.
- Metaphor only makes sense if it is a part of a process. A process of learning. It doesn’t do much good on its own.
## How metaphors work in helping us understand things
Teachers love explaining things. Students love understanding things. On the rare occasions that the two coincide, the feeling of joy shines like a beacon for the power of explanation. Teachers tell stories of seeing the “lightbulbs come on” in their students' eyes. Students remember fondly the ecstatic moments of sudden illumination as their teacher’s words suddenly lit up the darkness within them. Thus the myth of teaching as explaining and learning as understanding those explanations was born.
Most of the more powerful explanations rely on metaphor in the broadest possible sense. In fact, all explanation is to some extent metaphorical in that it provides a projection from one domain of understanding onto another. Metaphor brings out the familiar - or ex plains it - in the unfamiliar. Or so the story goes.
We can think of metaphoric projection as putting two thin sheets of paper over each other and looking at them against a bright light. What can be on these sheets? Sketches, images, words or even just smudges of color. The projection then obscures certain things and shows others in new contexts. Sometimes, with more complex slides we may see completely new shapes and color hues. The process of making sense of the metaphor then involves slight adjustments in how those two sheets align against one another. This can be described as the metaphor giving a new structure to the target domain.
Another way to think about metaphoric projection is as two sets of items which are mapped onto each other. We can put the sets side by side and draw lines between items we think match. Or we can take them out and place them side by side in a new set. We often see them displayed in this way.
Note: This way of thinking about metaphor started with Lakoff and Johnson’s ‘Metaphors we live by’ from 1980. This led to the formulation of the Conceptual Metaphor Theory. It was later developed into a more general theory of frames or mental models by Turner and Fauconnier (2002) known as the theory of conceptual integration or blending. But it can also be found in Donald A Schön’s ‘Displacement of Concepts’ from 1963 which indirectly inspired Lakoff and Johnson.
But despite all this, it is easy to overlook that in order to form a projection from one mental space into another, we have to have some structure in both. In fact, metaphor often assumes equal knowledge of both domains (https://medium.com/metaphor-hacker/fruit-loops-and-metaphors-metaphors-are-not-about-explaining-the-abstract-through-concrete-but-4e0209dd70b6), and in the process of making a projection from one another, a new previously unimagined structure emerges that is a blend of both domains. Because of the complexity, it is hard to give brief examples, but Turner’s and Fauconnier’s ‘The Way We Think’ is full of very illuminating case studies.
But it is also not at all uncommon for metaphor to borrow from a domain we know much less about to elucidate a domain we know a lot about. For example, if I hear, ‘don’t go into that office, the boss is on a warpath’, I understand a lot more about the boss’s behaviour than I do about any warpaths. Here, only the general feeling of ferocity is transferred with none of the possible association of weaponry or military supply lines.
Metaphor is also always partial. It would make no sense to project every aspect of both domains onto one another. But the ability to understand which bits it makes sense to project and which must be left out also requires at least some understanding of both domains. To understand what we mean when we call a piece of software a ‘virus’ we must know enough about computers to know that the infection cannot be transmitted through simple touch.
Metaphor at its most powerful helps us understand both domains better. It also often results in the creation of new understanding of both domains as we strive to find the limits of possible cross-domain mappings. Often, this happens with honest historical explanations of the present. By comparing the Iraq war to Vietnam, we may only choose to transfer the feeling of emotion and loss associated with the former. But we may also choose to explore both in their own right to find the best way in which they project on to one another. And this gives us new understanding of both.
## Three uses of metaphor in explanation
There are many ways to classify the uses of metaphors, I’ve outlined some in an early paper (/2013/04/how-we-use-metaphors/). But for the purposes of metaphor in explanation, I’d like to offer three broad types: 1. Metaphor as invitation; 2. Metaphor as instrument; 3. Metaphor as catalyst. I fear that the first type may be most common while only the second two play any real role in building understanding. These three types could also be viewed as forming a sort of process but this is not inherent in the definition.
As we will see, sometimes the same metaphor can serve all three roles, providing a certain thread through the process of learning. But most often, we need new metaphors for each type or stage.
### 1. Metaphor as invitation
Novice students often come to a new subject with no knowledge and a healthy dose of fear of the unknown. To help them feel more comfortable, teachers like to reach for metaphors relying on the familiar. This gives the learner a chance to grasp onto something while they build up sufficient mental representations of the new domain.
But this use of metaphor usually does not help understanding. It just provides emotional support along the arduous journey towards that understanding. It can also backfire. Teachers often spin up these kinds of metaphors in such a way that they assume an understanding of the unfamiliar. And it is only once students have bootstrapped themselves into some understanding of the subject that the metaphor starts to make any sense to them.
For instance (to use a famous example), we can teach students that the electrical current is like a flow of water. This certainly takes some fear out of the invisible world of electrons. But unless students have at least some prior understanding of electricity, they may ask questions like ‘how do you get the water into the wires?’
This type of metaphor can only be used for a fleeting moment and it must be followed by hard work of accumulating understanding of the new domain on its own terms. Perhaps with the use of more metaphors, this time of the instrumental kind.
### 2. Metaphor as an instrument
The instrumental use of metaphor for explanation is where real understanding starts to happen. But not all teachers are as good at it. In this case, the metaphor provides a way for the student to grasp the new subject. A lens to see it through, or a mental instrument to manipulate it with. Such metaphors are essential to the learning process. However, they do not rely on the moment of instant insight, which they can sometimes trigger, but rather on continuing exploration of the projection between the two domains. Their usefulness is less in the feeling of illumination than in their availability to be used over and over again.
For instance, electrical engineers may be able to make better judgments about certain properties of electrical circuits when they think of electrons as a flow of water. But in other instances, they may be better off when they think about electrons as lots of tiny balls rubbing against one another, generating heat. This metaphor can come up over and over to help them mentally manipulate the two domains.
Here, as with all metaphors, it is essential that we know when to let go. Or even better, when to switch to a different or even a contradictory metaphor. These instrumental metaphors can be local or global but it is rare that one will be enough.
### 3. Metaphor as catalyst
In the third use, the metaphor plays the role of a catalyst. Like a powder dissolved into a liquid, it makes a new substance in which both domains are transformed into one unified understanding. This is when the student transforms into a scholar. Making independent judgments, challenging the teacher’s own understanding, and ultimately becoming her own teacher. To work as a catalyst, the metaphor may be very rich and detailed or just a quick sketch resulting in a slight shift of perspective. But it always requires solid knowledge of the target domain.
Let’s continue with our electrical current example. Here, the student comes not only to understand that sometimes electricity behaves like a liquid and sometimes like a collection of particles, they also come to see the complexity of liquids and particles. They start making predictions both ways and ask questions like ‘What if we thought of the flow of water as a collection of particles?’, etc.
Here the metaphor becomes a process without an end. It spurs new mixtures and remixtures as one finds out more about the two (and often more) domains. Unlike with instrumental and invitational metaphors, it is no longer important that the metaphor be apt. It is just important that it is useful for new understandings or the possibilities of these new understandings. Donald Schön called one subtype of these ‘generative metaphor (/2011/06/language-learning-in-literature-as-a-source-domain-for-generative-metaphors-about-anything/)’.
But as with the other types, it is important that these metaphors come with some sort self-destruct mechanism.
What often happens is that these metaphors are taken up by those who presume that they map fully onto the target domain and that no other understanding of the target domain is necessary. I described how this is a problem with Schroedinger’s cat, or Lorenz’s hurricane-triggering butterflies (/2018/11/cats-and-butterflies-2-misunderstood-analogies-in-scientistic-discourse/).
What’s even worse, teachers often use these metaphors far too soon. This either confuses students or, worse, it gives them an illusion of understanding that they do not possess.
## How NOT to use metaphor to explain something - two case studies
### Case study 1: Metaphor gap in data science
My first case study of a bad use of explanation with metaphor is the podcast Data Skeptic (https://dataskeptic.com). In fact, listening to the most recent episodes prompted me to write this in the first place.
I must preface this by saying that I like the podcast and recommend it to others who want to understand modern data science. It covers important subjects and there is much to learn from it. Its one unfortunate feature, however, are certain episodes when the host, data scientist Kyle Polich, uses his wife, project manager and English major, Linh Da Tran as co-host and tries to explain concepts from abstract computational theory to her. Or rather at her.
This almost invariably fails. Not because Linh Da does not possess the raw intelligence or aptitude to understand these concepts but because Kyle confuses metaphor with explanation and explanation with understanding.
In two recent episodes, he attempted to explain attention in neural networks (https://dataskeptic.com/blog/episodes/2019/attention) and Neural Turing Machines (https://dataskeptic.com/blog/episodes/2019/neural-turing-machines). It was an unmitigated disaster. As the metaphors kept piling up, Linh Da finally cried out “I don’t know what you want me to understand”. That’s exactly the problem with a metaphor that only relies on the understanding of the source domain. It serves as a good invitation to the subject but as a very bad instrument for developing an understanding.
There are several problems with this set up that make it a bad place for too many metaphors. First, Linh Da is clearly just humoring Kyle. She’s vaguely interested in machine learning as a phenomenon but has no real interest in putting much work in to learning about how it works. This forces Kyle into more and more metaphors involving their pet bird Yoshi. These are useful socially and emotionally because they allow Linh Da to contribute to the discussion. But her contributions at every turn show that she cannot use any of the analogies to make useful inferences about the subject. She almost never brings up previous subjects. At the end of the episode on Neural Turing Machines, she asked who owns the Turing Machine. In all the torrent of analogies, Kyle neglected to stress that the Turing Machine is itself a metaphor. This is despite a prior episode where another guest explained why Turing Machines are important very clearly.
The conceit of the episodes is that data science can be explained even to English majors. That is certainly correct. But those majors must be willing to put in some work between episodes or have some prior knowledge. And as the subjects get more technical or abstract, the explanations have to get longer and include some practice time. And the amount of this practice needs to increase as if the practice was filling a funnel and not a test tube. Namely, to get from level B to C requires more work than getting from level A to B. Otherwise, the metaphors have nothing to hold on to. They constantly invite the student in but then offer no tools for going further. At best, they will confuse the learner and at worst, they will give them an illusion of understanding. About as useful as a seat belt made from masking tape.
While it is pleasant learning about these concepts through listening in on a married couple having a light-hearted conversation, at a certain stage, this pedagogic device just gets in the way of learning by the audience. Initially, the listener can just do their own metaphor mapping and ask the right questions in their head. But as the abstractness level increases, the host doing the explaining cannot go into sufficient depth because the co-host can’t keep up. And the increasingly convoluted and unnecessary metaphors just create a mental fog that descends over all.
I was particularly disappointed in the episode on Attention in neural networks which is something I wanted to learn more about. I found the initial metaphor of attention as a sort of memory span very useful but then it got stuck because Linh Da could not use it to go any further. This was because she was not given a chance to integrate the previous episodes where similar things were discussed. It was still useful to me because then I could go read about attention with a renewed perspective. But an opportunity for a deeper exploration of the metaphor was wasted.
This would have been fine if the episode was aimed at general public with no other understanding rather than an interested audience with some prior background. But even then, the general public would have needed more and different information to make any sense of it.
At one point Kyle, raised the possibility that maybe he wasn’t an effective teacher because Linh Da could not understand something he had explained. But in fact, he was not being a teacher at all. In this setting, he’s just a provider of images. Like a documentary from the Serengeti where the audience remembers there are lions, but could not place it on a map.
I can imagine that Kyle would be a very effective teacher with students who are interested in the subject and if he had a chance to take them through it step by step. And his use of metaphors would be a valuable contribution to that. But in the podcast, he’s only playing at being a teacher with Linh Da and she’s only pretending being a student. His only goals are getting her to answer questions within his metaphor that seem like she achieved comprehension. This means she never gets a chance to try out the structures of the source domain on the target domain. And because of this she never gets to develop any understanding that could later be used as a foundation for further metaphors. Without this, adding more to the mix feels like an avalanche of analogies.
### Case study 2: The explanation illusion at Wired magazine
But Data Skeptic is not the worst example of this type of pseudo-teaching by explanation. Only the most recent in my mind. A possibly much worse example is the Wired magazine series in which one expert supposedly explains a technical concept at 5 levels of difficulty: 5-7-year old, young teen, college student, graduate student, and another expert. These explanations often involve some level of metaphor, but they are mostly pointless. The conceit is that anybody can understand these concepts at “some level”. But the explanations do not equal understanding as is amply demonstrated in the videos. The people being explained to do not usually develop any new understanding. And it is doubtful whether the people watching do either.
Some of these are because the topic just is not appropriate to be explained to a certain audience. A 5 or 13-year old do not need to understand (nor do they have the background to) things like CRISPR or the Conectome. At best, they may learn which discipline they belong to, but that’s just teaching them a new name. No understanding of the phenomena is necessary.
But even when the understanding is well within reach and might have its use, the ‘expert’ fumbles. Thus the great and inventive musician Jacob Collier failed to explain the concept of ‘harmony’ to any of his charges. First, he tried to convince a five year old that harmony is a way of expressing a feeling with music (as opposed to melody). This is not only too abstract, it is also wrong. Both harmony and melody express feelings. But harmony is different notes played on top of one another rather than in sequence as in a melody (the feelings come from the pitch distance between the tones). This is well within the scope of understanding of a 5-year old when accompanied by some examples. No elaborate metaphors are necessary. But Jacob Collier goes into a very abstract explanation concluding with the most pointless question in any teacher’s arsenal: ‘does this make sense?’ to which he gets a an ‘uhuh’ from the child who clearly has no clue.
But explaining anything to 5-year-olds is hard. So does he do better with a teen? No. He still sticks with the metaphor of harmony as adding emotion to a melody. But then he mixes in the idea of harmony being a journey. To illustrate this, he goes from demonstrating a simple major / minor cord distinction to a jazz chord substitution. Which is wonderful and impresses the student but does not illustrate the concept of harmony to her.
No explanation happens at the higher levels either because all of the others (culminating in jazz giant Herbie Hancock) know the key concepts. So Collier just chats with them about harmonization and reharmonization. Which also reveals that that’s what he had in his mind with the 5-year old and the teen - he was just explaining a much more advanced concept under the label of the simple one.
One of the commenters on the video made an astute observation:
“it’s interesting how in the earlier levels it has to do more with theory and as you get higher up the level it goes back to nature and life experience and emotions. It’s almost as if, as the complexity increases, there’s also a level of fundamental basic understanding of nature and how it goes hand in hand at the most complex level” (emusik97531 [DL fixed small typos])
Essentially, as the level of the underlying understanding grows, the simple metaphor of journey, place and feeling have the most impact. At the lower levels, they just hang in there, not doing much of anything. They may feel like an invitation, but they don’t have any way to be used as a tool for understanding.
At the higher levels, Collier also shows that maybe he could be a great teacher to somebody closer to his level of skill and understanding. But it also reveals the pointlessness of an isolated act of explanation with (or without) metaphor if it is not supported by the hard work of making the connections necessary for the metaphor to become a proper instrument or a catalyst.
This is not a particular critique of Jacob Collier who is a great teacher to students at Berklee but rather of the whole set up of the series by Wired. Nobody could succeed in this setting. The concept is either going to be hard at the low levels or too basic at the higher ones.
## The inglorious history of metaphorical explanation in education
Collier and Polich, as well as countless others, are in illustrious company of people who overestimate what explanation can do in the process of learning.
Socrates in a famous scene from ‘Meno’ (https: