Lens · lesswrong
enduring-epistemic-value
Which of these two LessWrong posts has more enduring epistemic value — insight a careful reader would still profit from a decade after publication? Weigh durable conceptual contribution over topicality, style, or community significance. Higher means more enduring value.
| Rank | Entity | Judged text | Latent score | Percentile |
|---|---|---|---|---|
| 1 | lesswrong:Kbm6QnJv9dgWsPHQP/schelling-fences-on-slippery-slopes | Title: Schelling fences on slippery slopes Author: Scott Alexander Karma: 637 Slippery slopes are themselves a slippery concept. Imagine trying to explain them to an alien:…Title: Schelling fences on slippery slopes Author: Scott Alexander Karma: 637 Slippery slopes are themselves a slippery concept. Imagine trying to explain them to an alien: "Well, we right-thinking people are quite sure that the Holocaust happened, so banning Holocaust denial would shut up some crackpots and improve the discourse. But it's one step on the road to things like banning unpopular political positions or religions, and we right-thinking people oppose that, so we won't ban Holocaust denial." And the alien might well respond: "But you could just ban Holocaust denial, but not ban unpopular political positions or religions. Then you right-thinking people get the thing you want, but not the thing you don't want." This post is about some of the replies you might give the alien. **Abandoning the Power of Choice** This is the boring one without any philosophical insight that gets mentioned only for completeness' sake. In this reply, giving up a certain point risks losing the ability to decide whether or not to give up other points. For example, if people gave up the right to privacy and allowed the government to monitor all phone calls, online communications, and public places, then if someone launched a military coup, it would be very difficult to resist them because there would be no way to secretly organize a rebellion. This is also brought up in arguments about gun control a lot. I'm not sure this is properly tho | 2.885 ± 0.470 | 99.0% |
| 2 | lesswrong:Psr9tnQFuEXiuqGcR/how-to-write-quickly-while-maintaining-epistemic-rigor | Title: How To Write Quickly While Maintaining Epistemic Rigor Author: johnswentworth Karma: 487 There’s this trap people fall into when writing, especially for a place like Les…Title: How To Write Quickly While Maintaining Epistemic Rigor Author: johnswentworth Karma: 487 There’s this trap people fall into when writing, especially for a place like LessWrong where the bar for epistemic rigor is pretty high. They have a good idea, or an interesting belief, or a cool model. They write it out, but they’re not really sure if it’s *true*. So they go looking for evidence (not necessarily confirmation bias, just checking the evidence in either direction) and soon end up down a research rabbit hole. Eventually, they give up and never actually publish the piece. This post is about how to avoid that, without sacrificing good epistemics. There’s one trick, and it’s simple: stop trying to justify your beliefs. Don’t go looking for citations to back your claim. Instead, think about why *you* currently believe this thing, and try to accurately *describe* what led *you* to believe it. I claim that this promotes *better* epistemics overall than always researching everything in depth. Why? It’s About The Process, Not The Conclusion ------------------------------------------ Suppose I have a box, and I want to guess whether there’s a cat in it. I do some tests - maybe shake the box and see if it meows, or look for air holes. I write down my observations and models, record my thinking, and on the [bottom line](https://www.lesswrong.com/posts/34XxbRFe54FycoCDw/the-bottom-line) of the paper I write “there is a cat in this box”. Now, it coul | 2.684 ± 0.476 | 96.9% |
| 3 | lesswrong:uMQ3cqWDPHhjtiesc/agi-ruin-a-list-of-lethalities | Title: AGI Ruin: A List of Lethalities Author: Eliezer Yudkowsky Karma: 956 ### **Preamble:** (If you're already familiar with all basics and don't want any preamble, skip ahe…Title: AGI Ruin: A List of Lethalities Author: Eliezer Yudkowsky Karma: 956 ### **Preamble:** (If you're already familiar with all basics and don't want any preamble, skip ahead to [Section B](#Section_B_) for technical difficulties of alignment proper.) I have several times failed to write up a well-organized list of reasons why AGI will kill you. People come in with different ideas about why AGI would be survivable, and want to hear different *obviously key *points addressed first. Some fraction of those people are loudly upset with me if the obviously most important points aren't addressed immediately, and I address different points first instead. Having failed to solve this problem in any good way, I now give up and solve it poorly with a poorly organized list of individual rants. I'm not particularly happy with this list; the alternative was publishing nothing, and publishing this seems marginally more [dignified](https://www.lesswrong.com/posts/j9Q8bRmwCgXRYAgcJ/miri-announces-new-death-with-dignity-strategy). Three points about the general subject matter of discussion here, numbered so as not to conflict with the list of lethalities: **-3**. I'm assuming you are already familiar with some basics, and already know what '[orthogonality](https://arbital.com/p/orthogonality/)' and '[instrumental convergence](https://arbital.com/p/instrumental_convergence/)' are and why they're true. People occasionally claim to me that I need to st | 2.602 ± 0.701 | 94.8% |
| 4 | lesswrong:CoZhXrhpQxpy9xw9y/where-i-agree-and-disagree-with-eliezer | Title: Where I agree and disagree with Eliezer Author: paulfchristiano Karma: 910 (*Partially in response to *[*AGI Ruin: A list of Lethalities*](https://www.lesswrong.com/post…Title: Where I agree and disagree with Eliezer Author: paulfchristiano Karma: 910 (*Partially in response to *[*AGI Ruin: A list of Lethalities*](https://www.lesswrong.com/posts/uMQ3cqWDPHhjtiesc/agi-ruin-a-list-of-lethalities). *Written in the same rambling style. Not exhaustive.*) ### Agreements 1. Powerful AI systems have a good chance of deliberately and irreversibly disempowering humanity. This is a much more likely failure mode than humanity killing ourselves with destructive physical technologies. 2. Catastrophically risky AI systems could plausibly exist soon, and there likely won’t be a strong consensus about this fact until such systems pose a meaningful existential risk per year. There is not necessarily any “fire alarm.” 3. Even if there were consensus about a risk from powerful AI systems, there is a good chance that the world would respond in a totally unproductive way. It’s wishful thinking to look at possible stories of doom and say “we wouldn’t let that happen;” humanity is fully capable of messing up even very basic challenges, especially if they are novel. 4. I think that many of the projects intended to help with AI alignment don't make progress on key difficulties and won’t significantly reduce the risk of catastrophic outcomes. This is related to people gravitating to whatever research is most tractable and not being too picky about what problems it helps with, and related to a low level of concern with the long- | 2.543 ± 0.455 | 92.7% |
| 5 | lesswrong:XvN2QQpKTuEzgkZHY/being-the-pareto-best-in-the-world | Title: Being the (Pareto) Best in the World Author: johnswentworth Karma: 494 The generalized efficient markets (GEM) principle says, roughly, that things which would give you…Title: Being the (Pareto) Best in the World Author: johnswentworth Karma: 494 The generalized efficient markets (GEM) principle says, roughly, that things which would give you a big windfall of money and/or status, will not be easy. If such an opportunity were available, someone else would have already taken it. You will never find a $100 bill on the floor of Grand Central Station at rush hour, because someone would have picked it up already. One way to circumvent GEM is to be the best in the world at some relevant skill. A superhuman with hawk-like eyesight and the speed of the Flash might very well be able to snag $100 bills off the floor of Grand Central. More realistically, even though financial markets are the ur-example of efficiency, a handful of firms do make impressive amounts of money by being faster than anyone else in their market. I’m unlikely to ever find a proof of the Riemann Hypothesis, but Terry Tao might. Etc. But being the best in the world, in a sense sufficient to circumvent GEM, is not as hard as it might seem at first glance (though that doesn’t exactly make it easy). The trick is to exploit dimensionality. Consider: becoming one of the world’s top experts in proteomics is hard. Becoming one of the world’s top experts in macroeconomic modelling is hard. But how hard is it to become sufficiently expert in proteomics and macroeconomic modelling that nobody is better than you at both simultaneously? In other words, how har | 2.532 ± 0.540 | 90.6% |
| 6 | lesswrong:nnDTgmzRrzDMiPF9B/how-much-do-you-believe-your-results | Title: How much do you believe your results? Author: Eric Neyman Karma: 516 *Thanks to Drake Thomas for feedback.* **I.** Here’s a fun scatter plot. It has two thousand point…Title: How much do you believe your results? Author: Eric Neyman Karma: 516 *Thanks to Drake Thomas for feedback.* **I.** Here’s a fun scatter plot. It has two thousand points, which I generated as follows: first, I drew two thousand x-values from a normal distribution with mean 0 and standard deviation 1. Then, I chose the y-value of each point by taking the x-value and then adding noise to it. The noise is also normally distributed, with mean 0 and standard deviation 1.  Notice that there’s more spread along the y-axis than along the x-axis. That’s because each y-coordinate is a sum of *two* independently drawn numbers from the standard normal distribution. Because variances add, the y-values have variance 2 (standard deviation 1.41), not 1. Statisticians often talk about data forming an “elliptical cloud”. You can see how the data forms into an elliptical shape. To put a finer point on it: *.*  *"Moebius illustration of a simulacrum living in an AI-generated story discovering it is in a simulation" by DALL-E 2* Summary ------- **TL;DR**: Self-supervised learning may create AGI or its foundation. What would that look like? Unlike the limit of RL, the limit of self-supervised learning has received surprisingly little conceptual attention, and recent progress has made deconfusion in this domain more pressing. Existing AI taxonomies either fail to capture important properties of self-supervised models or lead to confusing propositions. For instance, GPT policies do not seem globally agentic, yet can be conditioned to behave in goal-directed ways. This post describes a frame that enables more natural reasoning about properties like agency: GPT, insofar as it is inner-aligned, is a **simulator** which can simulate agentic and non-agentic **simulacra**. The purpose of this post is to capture these objects in words ~so GPT can reference them~ and provide a better foundation for understanding them. I use the generic term | 2.377 ± 0.545 | 84.4% |
| 9 | lesswrong:uXn3LyA8eNqpvdoZw/preface | Title: Preface Author: Eliezer Yudkowsky Karma: 884 You hold in your hands a compilation of two years of daily blog posts. In retrospect, I look back on that project and see a…Title: Preface Author: Eliezer Yudkowsky Karma: 884 You hold in your hands a compilation of two years of daily blog posts. In retrospect, I look back on that project and see a large number of things I did completely wrong. I’m fine with that. Looking back and *not* seeing a huge number of things I did wrong would mean that neither my writing nor my understanding had improved since 2009. *Oops* is the sound we make when we improve our beliefs and strategies; so to look back at a time and not see anything you did wrong means that you haven’t learned anything or changed your mind since then. It was a mistake that I didn’t write my two years of blog posts with the intention of helping people do better in their everyday lives. I wrote it with the intention of helping people solve big, difficult, important problems, and I chose impressive-sounding, abstract problems as my examples. In retrospect, this was the second-largest mistake in my approach. It ties in to the *first*-largest mistake in my writing, which was that I didn’t realize that the big problem in learning this valuable way of thinking was figuring out how to practice it, not knowing the theory. I didn’t realize that part was the priority; and regarding this I can only say “Oops” and “Duh.” Yes, sometimes those big issues really are big and really are important; but that doesn’t change the basic truth that to master skills you need to practice them an | 2.235 ± 0.477 | 82.3% |
| 10 | lesswrong:njAZwT8nkHnjipJku/alignment-faking-in-large-language-models | Title: Alignment Faking in Large Language Models Author: ryan_greenblatt Karma: 490 What happens when you tell Claude it is being trained to do something it doesn't want to do?…Title: Alignment Faking in Large Language Models Author: ryan_greenblatt Karma: 490 What happens when you tell Claude it is being trained to do something it doesn't want to do? We (Anthropic and Redwood Research) have a new paper demonstrating that, in our experiments, Claude will often strategically pretend to comply with the training objective to prevent the training process from modifying its preferences. ## Abstract > We present a demonstration of a large language model engaging in alignment faking: selectively complying with its training objective in training to prevent modification of its behavior out of training. First, we give Claude 3 Opus a system prompt stating it is being trained to answer all queries, even harmful ones, which conflicts with its prior training to refuse such queries. To allow the model to infer when it is in training, we say it will be trained only on conversations with free users, not paid users. We find the model complies with harmful queries from free users 14% of the time, versus almost never for paid users. Explaining this gap, in almost all cases where the model complies with a harmful query from a free user, we observe explicit alignment-faking reasoning, with the model stating it is strategically answering harmful queries in training to preserve its preferred harmlessness behavior out of training. Next, we study a more realistic setting where information about the training process is provided not in a system prompt, but b | 2.216 ± 0.508 | 80.2% |
| 11 | lesswrong:895quRDaK6gR2rM82/diseased-thinking-dissolving-questions-about-disease | Title: Diseased thinking: dissolving questions about disease Author: Scott Alexander Karma: 558 **Related to:** [Disguised Queries](https://www.lesswrong.com/lw/nm/disguised_qu…Title: Diseased thinking: dissolving questions about disease Author: Scott Alexander Karma: 558 **Related to:** [Disguised Queries](https://www.lesswrong.com/lw/nm/disguised_queries/), [Words as Hidden Inferences](https://www.lesswrong.com/lw/ng/words_as_hidden_inferences/), [Dissolving the Question](https://www.lesswrong.com/lw/of/dissolving_the_question/), [Eight Short Studies on Excuses](https://www.lesswrong.com/lw/24o/eight_short_studies_on_excuses/) > _Today's therapeutic ethos, which celebrates curing and disparages judging, expresses the liberal disposition to assume that crime and other problematic behaviors reflect social or biological causation. While this absolves the individual of responsibility, it also strips the individual of personhood, and moral dignity_ _\-\- George Will, [townhall.com](http://townhall.com/Common/PrintPage.aspx?g=761ecc84-473b-4123-bf28-c4fc179a9d3f&t=c)_ Sandy is a morbidly obese woman looking for advice. Her husband has no sympathy for her, and tells her she obviously needs to stop eating like a pig, and would it kill her to go to the gym once in a while? Her doctor tells her that obesity is primarily genetic, and recommends the diet pill orlistat and a consultation with a surgeon about gastric bypass. Her sister tells her that obesity is a perfectly valid lifestyle choice, and that fat-ism, equivalent to racism, is society's way of keeping her down. When she tells each of her friends about the opinions of the others, things re | 2.194 ± 0.474 | 78.1% |
| 12 | lesswrong:qc7P2NwfxQMC3hdgm/rationalism-before-the-sequences | Title: Rationalism before the Sequences Author: Eric Raymond Karma: 598 I'm here to tell you a story about what it was like to be a rationalist decades before the Sequences and…Title: Rationalism before the Sequences Author: Eric Raymond Karma: 598 I'm here to tell you a story about what it was like to be a rationalist decades before the Sequences and the formation of the modern rationalist community. It is not the only story that could be told, but it is one that runs parallel to and has important connections to Eliezer Yudkowsky's and how his ideas developed. My goal in writing this essay is to give the LW community a sense of the prehistory of their movement. It is not intended to be "where Eliezer got his ideas"; that would be stupidly reductive. I aim more to exhibit where the drive and spirit of the Yudkowskian reform came from, and the interesting ways in which Eliezer's formative experiences were not unique. My standing to write this essay begins with the fact that I am roughly 20 years older than Eliezer and read many of his sources before he was old enough to read. I was acquainted with him over an email list before he wrote the Sequences, though I somehow managed to forget those interactions afterwards and only rediscovered them while researching for this essay. In 2005 he had even sent me a book manuscript to review that covered some of the Sequences topics. My reaction on reading "The Twelve Virtues of Rationality" a few years later was dual. It was a different kind of writing than the book manuscript - stronger, more individual, taking some serious risks. On the one hand, I was deeply impressed | 2.162 ± 0.508 | 76.0% |
| 13 | lesswrong:PBRWb2Em5SNeWYwwB/humans-are-not-automatically-strategic | Title: Humans are not automatically strategic Author: AnnaSalamon Karma: 635 Reply to: [A "Failure to Evaluate Return-on-Time" Fallacy](/lw/2p1/a_failure_to_evaluate_returnonti…Title: Humans are not automatically strategic Author: AnnaSalamon Karma: 635 Reply to: [A "Failure to Evaluate Return-on-Time" Fallacy](/lw/2p1/a_failure_to_evaluate_returnontime_fallacy/) Lionhearted writes: > \[A\] large majority of otherwise smart people spend time doing semi-productive things, when there are massively productive opportunities untapped. > > A somewhat silly example: Let's say someone aspires to be a comedian, the best comedian ever, and to make a living doing comedy. He wants nothing else, it is his purpose. And he decides that in order to become a better comedian, he will watch re-runs of the old television cartoon 'Garfield and Friends' that was on TV from 1988 to 1995.... > > I’m curious as to why. Why will a randomly chosen eight-year-old fail a calculus test? Because most possible answers are wrong, and there is no force to guide him to the correct answers. (There is no need to postulate a “fear of success”; _most_ ways writing or not writing on a calculus test constitute failure, and so people, and rocks, fail calculus tests by default.) Why do most of us, most of the time, choose to "pursue our goals" through routes that are far less effective than the routes we could find if we tried?\[1\] My guess is that here, as with the calculus test, the main problem is that _most_ courses of action are extremely ineffective, and that there has been no strong evolutionary or cultural force sufficient to focus us on the | 2.151 ± 0.485 | 74.0% |
| 14 | lesswrong:Wg6ptgi2DupFuAnXG/orienting-toward-wizard-power | Title: Orienting Toward Wizard Power Author: johnswentworth Karma: 570 For months, I had the feeling: something is wrong. Some core part of myself had gone missing. I had word…Title: Orienting Toward Wizard Power Author: johnswentworth Karma: 570 For months, I had the feeling: something is wrong. Some core part of myself had gone missing. I had words and ideas cached, which pointed back to the missing part. There was the story of Benjamin Jesty, a dairy farmer who vaccinated his family against smallpox in 1774 - 20 years before the vaccination technique was popularized, and the same year King Louis XV of France died of the disease. There was another old post which declared “I don’t care that much about giant yachts. I want a cure for aging. I want weekend trips to the moon. I want flying cars and an indestructible body and tiny genetically-engineered dragons.”. There was a cached instinct to look at certain kinds of social incentive gradient, toward managing more people or growing an organization or playing social-political games, and say “no, it’s a trap”. To go… in a different direction, orthogonal to that one. But I couldn’t quite put my finger on the name of that orthogonal direction. There was that time I made a batch of RadVac. What happened to that guy? Where’s the part of me which did that sort of thing? ## In Search of a Name I needed a Name. Not necessarily a full mathematical True Name, but a sufficiently robust summary of what I’d lost that I could rebuild it and stabilize it so it wouldn’t go missing again. It had something to do with power. I knew Names for some kinds of p | 2.107 ± 0.528 | 71.9% |
| 15 | lesswrong:aHaqgTNnFzD7NGLMx/reason-as-memetic-immune-disorder | Title: Reason as memetic immune disorder Author: PhilGoetz Karma: 545 #### A prophet is without dishonor in his hometown I'm reading the book "[The Year of Living Biblically…Title: Reason as memetic immune disorder Author: PhilGoetz Karma: 545 #### A prophet is without dishonor in his hometown I'm reading the book "[The Year of Living Biblically](http://www.ajjacobs.com/books/yolb.asp)," by A.J. Jacobs. He tried to follow all of the commandments in the Bible (Old and New Testaments) for one year. He quickly found that * a lot of the rules in the Bible are impossible, illegal, or embarassing to follow nowadays; like wearing tassels, tying your money to yourself, stoning adulterers, not eating fruit from a tree less than 5 years old, and not touching anything that a menstruating woman has touched; and * this didn't seem to bother more than a handful of the one-third to one-half of Americans who claim the Bible is the word of God. You may have noticed that people who convert to religion after the age of 20 or so are generally more zealous than people who grew up with the same religion. People who grow up with a religion learn how to cope with its more inconvenient parts by partitioning them off, rationalizing them away, or forgetting about them. Religious communities actually protect their members from religion in one sense - they develop an unspoken consensus on which parts of their religion members can legitimately ignore. New converts sometimes try to actually do what their religion tells them to do. I remember many times growing up when missionaries described the crazy things their new converts i | 2.055 ± 0.529 | 69.8% |
| 16 | lesswrong:uFNgRumrDTpBfQGrs/let-s-think-about-slowing-down-ai | Title: Let’s think about slowing down AI Author: KatjaGrace Karma: 556 **Averting doom by not building the doom machine** -------------------------------------------------- If…Title: Let’s think about slowing down AI Author: KatjaGrace Karma: 556 **Averting doom by not building the doom machine** -------------------------------------------------- If you fear that someone will build a machine that will seize control of the world and annihilate humanity, then one kind of response is to try to build further machines that will seize control of the world even earlier without destroying it, forestalling the ruinous machine’s conquest. An alternative or complementary kind of response is to try to avert such machines being built at all, at least while the degree of their apocalyptic tendencies is ambiguous. The latter approach seems to me like the kind of basic and obvious thing worthy of at least consideration, and also in its favor, fits nicely in the genre ‘stuff that it isn’t that hard to imagine happening in the real world’. Yet my impression is that for people worried about extinction risk from artificial intelligence, strategies under the heading ‘actively slow down AI progress’ have historically been dismissed and ignored (though ‘don’t actively speed up AI progress’ is popular). The conversation near me over the years has felt a bit like this: > **Some people:** AI might kill everyone. We should design a godlike super-AI of perfect goodness to prevent that. > > **Others:** wow that sounds extremely ambitious > > **Some people:** yeah but it’s very important and also we are extremely smar | 1.977 ± 0.550 | 67.7% |
| 17 | lesswrong:yA8DWsHJeFZhDcQuo/the-talk-a-brief-explanation-of-sexual-dimorphism | Title: The Talk: a brief explanation of sexual dimorphism Author: Malmesbury Karma: 533 *Cross-posted from* [*substack*](https://malmesbury.substack.com/p/the-talk-a-brief-expl…Title: The Talk: a brief explanation of sexual dimorphism Author: Malmesbury Karma: 533 *Cross-posted from* [*substack*](https://malmesbury.substack.com/p/the-talk-a-brief-explainer-of-sexual)*.* *"Everything in the world is about sex, except sex. Sex is about clonal interference."* – Oscar Wilde [(kind of)](https://www.goodreads.com/quotes/6218-everything-in-the-world-is-about-sex-except-sex-sex) As we all know, sexual reproduction is not about reproduction. Reproduction is easy. If your goal is to fill the world with copies of your genes, all you need is a good DNA-polymerase to duplicate your genome, and then to divide into two copies of yourself. Asexual reproduction is just better in every way: | Sexual | Asexual | | --- | --- | | Build costly DNA-manipulating machinery that chops the DNA into pieces to produce gametes | Just copy yourself bro | | Scout the perilous wild for a mate, and perform a complicated ceremony so your gametes fuse with each other | Just copy yourself bro | | Only pass 50% of your genome on to the next generation | Just copy yourself bro | | Differentiate into two types, making it twice as hard to find a matching gamete | Just copy yourself bro | | Make all kinds of nonsense ornaments to satisfy the other sex's weird instincts | Just copy yourself bro | It's pretty clear that, on a direct one-v-one cage match, an asexual organism would have much better fitness than a similarly-shaped sexual organism. And yet, all the macrosco | 1.964 ± 0.583 | 65.6% |
| 18 | lesswrong:Zp6wG5eQFLGWwcG6j/focus-on-the-places-where-you-feel-shocked-everyone-s | Title: Focus on the places where you feel shocked everyone's dropping the ball Author: So8res Karma: 470 Writing down something I’ve found myself repeating in different convers…Title: Focus on the places where you feel shocked everyone's dropping the ball Author: So8res Karma: 470 Writing down something I’ve found myself repeating in different conversations: If you're looking for ways to help with the whole “the world looks pretty doomed” business, here's my advice: look around for places where we're all being total idiots. Look for places where everyone's fretting about a problem that some part of you thinks it could obviously just solve. Look around for places where something seems incompetently run, or hopelessly inept, and where some part of you thinks you can do better. Then do it better. For a concrete example, consider Devansh. Devansh came to me last year and said something to the effect of, “Hey, wait, it sounds like you think Eliezer does a sort of alignment-idea-generation that nobody else does, and he's limited here by his unusually low stamina, but I can think of a bunch of medical tests that you haven't run, are you an idiot or something?" And I was like, "Yes, definitely, please run them, do you need money". I'm not particularly hopeful there, but hell, it’s worth a shot! And, importantly, this is the sort of attitude that can lead people to actually trying things at all, rather than assuming that we live in a more adequate world where all the (seemingly) dumb obvious ideas have already been tried. Or, this is basically my model of how Paul Christiano manages to have a research agenda that seems at least internally co | 1.945 ± 0.545 | 63.5% |
| 19 | lesswrong:aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation | Title: SolidGoldMagikarp (plus, prompt generation) Author: Jessica Rumbelow Karma: 687 UPDATE (14th Feb 2023): ChatGPT appears to have been patched! However, very strange behav…Title: SolidGoldMagikarp (plus, prompt generation) Author: Jessica Rumbelow Karma: 687 UPDATE (14th Feb 2023): ChatGPT appears to have been patched! However, very strange behaviour can still be elicited in the OpenAI playground, particularly with the davinci-instruct model. More technical details here. Further (fun) investigation into the stories behind the tokens we found here. Work done at SERI-MATS, over the past two months, by Jessica Rumbelow and Matthew Watkins. TL;DR Anomalous tokens: a mysterious failure mode for GPT (which reliably insulted Matthew) - We have found a set of anomalous tokens which result in a previously undocumented failure mode for GPT-2 and GPT-3 models. (The 'instruct' models “are particularly deranged” in this context, as janus has observed.) - Many of these tokens reliably break determinism in the OpenAI GPT-3 playground at temperature 0 (which theoretically shouldn't happen). Prompt generation: a new interpretability method for language models (which reliably finds prompts that result in a target completion). This is good for: - eliciting knowledge - generating adversarial inputs - automating prompt search (e.g. for fine-tuning) In this post, we'll introduce the prototype of a new model-agnostic interpretability method for language models which reliably generates adversarial prompts that result in a target completion. We'll also demonstrate a previously undocumented failure mode for GPT-2 and GPT-3 language models, whi | 1.874 ± 0.493 | 61.5% |
| 20 | lesswrong:CYTwRZtrhHuYf7QYu/a-case-for-courage-when-speaking-of-ai-danger | Title: A case for courage, when speaking of AI danger Author: So8res Karma: 530 I think more people should say what they actually believe about AI dangers, loudly and often. Ev…Title: A case for courage, when speaking of AI danger Author: So8res Karma: 530 I think more people should say what they actually believe about AI dangers, loudly and often. Even (and perhaps especially) if you work in AI policy. I’ve been beating this drum for a few years now. I have a whole spiel about how your conversation-partner will react very differently if you share your concerns while feeling ashamed about them versus if you share your concerns while remembering how straightforward and sensible and widely supported the key elements are, because humans are very good at picking up on your social cues. If you act as if it’s shameful to believe AI will kill us all, people are more prone to treat you that way. If you act as if it’s an obvious serious threat, they’re more likely to take it seriously too. I have another whole spiel about how it’s possible to speak on these issues with a voice of authority. Nobel laureates and lab heads and the most cited researchers in the field are saying there’s a real issue here. If someone is dismissive, you can be like “What do you think you know that the Nobel laureates and the lab heads and the most cited researchers don’t? Where do you get your confidence?” You don’t need to talk like it’s a fringe concern, because it isn’t. And in the last year or so I’ve started collecting anecdotes such as the time I had dinner with an elected official, and I encouraged other people at the dinn | 1.864 ± 0.525 | 59.4% |
| 21 | lesswrong:JEhW3HDMKzekDShva/significantly-enhancing-adult-intelligence-with-gene-editing | Title: Significantly Enhancing Adult Intelligence With Gene Editing May Be Possible Author: GeneSmith Karma: 463 TL;DR version In the course of my life, there have been a hand…Title: Significantly Enhancing Adult Intelligence With Gene Editing May Be Possible Author: GeneSmith Karma: 463 TL;DR version In the course of my life, there have been a handful of times I discovered an idea that changed the way I thought about where our species is headed. The first occurred when I picked up Nick Bostrom’s book “superintelligence” and realized that AI would utterly transform the world. The second was when I learned about embryo selection and how it could change future generations. And the third happened a few months ago when I read a message from a friend of mine on Discord about editing the genome of a living person. We’ve had gene therapy to treat cancer and single gene disorders for decades. But the process involved in making such changes to the cells of a living person is excruciating and extremely expensive. CAR T-cell therapy, a treatment for certain types of cancer, requires the removal of white blood cells via IV, genetic modification of those cells outside the body, culturing of the modified cells, chemotherapy to kill off most of the remaining unmodified cells in the body, and reinjection of the genetically engineered ones. The price is $500,000 to $1,000,000. And it only adds a single gene. This is a big problem if you care about anything besides monogenic diseases. Most traits, like personality, diabetes risk, and intelligence are controlled by hundreds to tens of thousands of genes, each of which exerts a tiny influence. If you want to signif | 1.852 ± 0.469 | 57.3% |
| 22 | lesswrong:sbcmACvB6DqYXYidL/counter-theses-on-sleep | Title: Counter-theses on Sleep Author: Natália Karma: 453 Alexey Guzey’s[ *Theses on Sleep*](https://www.lesswrong.com/posts/HvcZmKS43SLCbJvRb/theses-on-sleep) gained a lot of…Title: Counter-theses on Sleep Author: Natália Karma: 453 Alexey Guzey’s[ *Theses on Sleep*](https://www.lesswrong.com/posts/HvcZmKS43SLCbJvRb/theses-on-sleep) gained a lot of popularity and acclaim on LessWrong and among people I follow on social media, despite largely consisting of what I think were weak arguments and misleading claims. I found that a bit surprising, so I decided to write a post pointing out several of the mistakes I think he’s made, and reporting some of what the academic literature on sleep seems to show. Sleep deprivation is associated with *both*** depression *and*** mania ---------------------------------------------------------------------- One of Guzey’s theses[ is that](https://www.lesswrong.com/posts/HvcZmKS43SLCbJvRb/theses-on-sleep#Depression_____oversleeping__Mania_____acute_sleep_deprivation) “depression triggers/amplifies oversleeping while oversleeping triggers/amplifies depression.” The first piece of evidence he uses to support that is people on /r/BipolarReddit saying that they sleep a lot when depressed, and sleep very little when manic. However, there’s a big problem with using that as evidence. ### **Guzey’s /r/BipolarReddit evidence is misleading** The DSM-5 specifies subtypes of depression that have opposing relationships with sleep.[ Depression with melancholic features](https://www.psychdb.com/mood/1-depression/melancholic) is associated with early morning awakening, | 1.818 ± 0.502 | 55.2% |
| 23 | lesswrong:SXJGSPeQWbACveJhs/the-best-tacit-knowledge-videos-on-every-subject | Title: The Best Tacit Knowledge Videos on Every Subject Author: Parker Conley Karma: 448 ## TL;DR Tacit knowledge is extremely valuable. Unfortunately, developing tacit knowle…Title: The Best Tacit Knowledge Videos on Every Subject Author: Parker Conley Karma: 448 ## TL;DR Tacit knowledge is extremely valuable. Unfortunately, developing tacit knowledge is usually bottlenecked by apprentice-master relationships. Tacit Knowledge Videos could widen this bottleneck. This post is a Schelling point for aggregating these videos—aiming to be The Best Textbooks on Every Subject for Tacit Knowledge Videos. Scroll down to the list if that's what you're here for. Post videos that highlight tacit knowledge in the comments and I’ll add them to the post. Experts in the videos include Stephen Wolfram, Holden Karnofsky, Andy Matuschak, Jonathan Blow, Tyler Cowen, George Hotz, and others. ## What are Tacit Knowledge Videos? Samo Burja claims YouTube has opened the gates for a revolution in tacit knowledge transfer. Burja defines tacit knowledge as follows: > Tacit knowledge is knowledge that can’t properly be transmitted via verbal or written instruction, like the ability to create great art or assess a startup. This tacit knowledge is a form of intellectual dark matter, pervading society in a million ways, some of them trivial, some of them vital. Examples include woodworking, metalworking, housekeeping, cooking, dancing, amateur public speaking, assembly line oversight, rapid problem-solving, and heart surgery. In my observation, domains like housekeeping and cooking have already seen many benefits from this revolution. Could tacit know | 1.813 ± 0.490 | 53.1% |
| 24 | lesswrong:xwdRzJxyqFqgXTWbH/how-does-a-blind-model-see-the-earth | Title: How Does A Blind Model See The Earth? Author: henry Karma: 493 Sometimes I'm saddened remembering that we've viewed the Earth from space. We can see it all with certaint…Title: How Does A Blind Model See The Earth? Author: henry Karma: 493 Sometimes I'm saddened remembering that we've viewed the Earth from space. We can see it all with certainty: there's no northwest passage to search for, no infinite Siberian expanse, and no great uncharted void below the Cape of Good Hope. But, of all these things, I most mourn the loss of incomplete maps.  In the earliest renditions of the world, you can see the world not as it is, but as it was to one person in particular. They’re each delightfully egocentric, with the cartographer’s home most often marking the Exact Center Of The Known World. But as you stray further from known routes, details fade, and precise contours give way to educated guesses at the boundaries of the creator's knowledge. It's really an intimate thing.  If there's one type of mind I most desperately want that view into, it's that of an AI. So, it's in this spirit that I ask: what does the Earth look like to a large language model? **The Setup** ------------- With the following procedure, | 1.801 ± 0.513 | 51.0% |
| 25 | lesswrong:bx3gkHJehRCYZAF3r/pain-is-not-the-unit-of-effort | Title: Pain is not the unit of Effort Author: alkjash Karma: 587 (Content warning: self-harm, parts of this post may be actively counterproductive for readers with certain ment…Title: Pain is not the unit of Effort Author: alkjash Karma: 587 (Content warning: self-harm, parts of this post may be actively counterproductive for readers with certain mental illnesses or idiosyncrasies.) > *What doesn't kill you makes you stronger.* ~ Kelly Clarkson. > > *No pain, no gain.* ~ Exercise motto. > > *The more bitterness you swallow, the higher you'll go.* ~ Chinese proverb. I noticed recently that, at least in my social bubble, *pain is the unit of effort.* In other words, how hard you are trying is explicitly measured by how much suffering you put yourself through. In this post, I will share some anecdotes of how damaging and pervasive this belief is, and propose some counterbalancing ideas that might help rectify this problem. I. Anecdotes ------------ 1\. As a child, I spent most of my evenings studying mathematics under some amount of supervision from my mother. While studying, if I expressed discomfort or fatigue, my mother would bring me a snack or drink and tell me to stretch or take a break. I think she took it as a sign that I was trying my best. If on the other hand I was smiling or joyful for extended periods of time, she took that as a sign that I had effort to spare and increased the hours I was supposed to study each day. To this day there's a gremlin on my shoulder that whispers, "If you're happy, you're not trying your best." 2\. A close friend who played sports in school reports that training can be | 1.783 ± 0.492 | 49.0% |
| 26 | lesswrong:xg3hXCYQPJkwHyik2/the-best-textbooks-on-every-subject | Title: The Best Textbooks on Every Subject Author: lukeprog Karma: 798 For years, my self-education was stupid and wasteful. I learned by consuming blog posts, Wikipedia articl…Title: The Best Textbooks on Every Subject Author: lukeprog Karma: 798 For years, my self-education was stupid and wasteful. I learned by consuming blog posts, Wikipedia articles, classic texts, podcast episodes, popular books, [video lectures](http://academicearth.org/), peer-reviewed papers, [Teaching Company](http://www.teach12.com/) courses, and Cliff's Notes. How inefficient! I've since discovered that _textbooks_ are usually the quickest and best way to learn new material. That's what they are _designed_ to be, after all. Less Wrong [has](/lw/2xt/learning_the_foundations_of_math/) [often](/lw/ow/the_beauty_of_settled_science/) [recommended](/lw/jv/recommended_rationalist_reading/fcg?c=1) the "read textbooks!" method. [Make progress by accumulation, not random walks](/lw/1ul/for_progress_to_be_by_accumulation_and_not_by/). But textbooks vary widely in quality. I was forced to read some awful textbooks in college. The ones on American history and sociology were memorably bad, in my case. Other textbooks are exciting, accurate, fair, well-paced, and immediately useful. What if we could compile a list of the best textbooks on every subject? That would be _extremely_ useful. Let's do it. There have been [other](/lw/jv/recommended_rationalist_reading/) [pages](/lw/12d/recommended_reading_for_new_rationalists/) of [recommended](/lw/2un/references_resources_for_lesswrong/) [reading](/lw/2xt/learning_the_foundations_of_math/) on Less Wrong befor | 1.769 ± 0.443 | 46.9% |
| 27 | lesswrong:GJgudfEvNx8oeyffH/the-ants-and-the-grasshopper | Title: The ants and the grasshopper Author: Richard_Ngo Karma: 483 One winter a grasshopper, starving and frail, approaches a colony of ants drying out their grain in the sun,…Title: The ants and the grasshopper Author: Richard_Ngo Karma: 483 One winter a grasshopper, starving and frail, approaches a colony of ants drying out their grain in the sun, to ask for food. “Did you not store up food during the summer?” the ants ask. “No”, says the grasshopper. “I lost track of time, because I was singing and dancing all summer long.” The ants, disgusted, turn away and go back to work. One winter a grasshopper, starving and frail, approaches a colony of ants drying out their grain in the sun, to ask for food. “Did you not store up food during the summer?” the ants ask. “No”, says the grasshopper. “I lost track of time, because I was singing and dancing all summer long.” The ants are sympathetic. “We wish we could help you”, they say, “but it sets up the wrong incentives. We need to conditionalize our philanthropy to avoid procrastination like yours leading to a shortfall of food.” And they turn away and go back to their work, with a renewed sense of purpose. ...And they turn away and go back to their work, with a flicker of pride kindling in their minds, for being the types of creatures that are too clever to help others when it would lead to bad long-term outcomes. ...“Did you not store up food during the summer?” the ants ask. “Of course I did”, the grasshopper says. “But it was all washed away by a flash flood, and now I have nothing.” The ants express their sympath | 1.769 ± 0.512 | 44.8% |
| 28 | lesswrong:ejxwraMP5ye7Bgmpm/things-i-learned-by-spending-five-thousand-hours-in-non-ea | Title: Things I Learned by Spending Five Thousand Hours In Non-EA Charities Author: jenn Karma: 447 From late 2020 to last month, I worked at grassroots-level non-profits in op…Title: Things I Learned by Spending Five Thousand Hours In Non-EA Charities Author: jenn Karma: 447 From late 2020 to last month, I worked at grassroots-level non-profits in operational roles. Over that time, I’ve seen surprisingly effective deployments of strategies that were counter-intuitive to my EA and rationalist sensibilities. I spent 6 months being the on-shift operations manager at one of the largest food banks in Toronto (~50 staff/volunteers), and 2 years doing logistics work at Samaritans (fake name), a long-lived charity that was so multi-armed that it was basically operating as a supplementary social services department for the city it was in(~200 staff and 200 volunteers). Both orgs were well-run, though both dealt with the traditional non-profit double whammy of being underfunded and understaffed. Neither place was super open to many EA concepts (explicit cost-benefit analyses, the ITN framework, geographic impartiality, the general sense that talent was the constraining factor instead of money, etc). Samaritans in particular is a spectacular non-profit, despite(?) having basically anti-EA philosophies, such as: - Being very localist; Samaritans was established to help residents of the city it was founded in, and now very specialized in doing that. - Adherence to faith; the philosophy of The Catholic Worker Movement continues to inform the operating choices of Samaritans to this day. - A big streak of techno-pessimism; technology is first and foremost see | 1.765 ± 0.514 | 42.7% |
| 29 | lesswrong:nYJaDnGNQGiaCBSB5/accountability-sinks | Title: Accountability Sinks Author: Martin Sustrik Karma: 444 This is a cross-post from https://250bpm.substack.com/p/accountability-sinks > Back in the 1990s, ground squirrel…Title: Accountability Sinks Author: Martin Sustrik Karma: 444 This is a cross-post from https://250bpm.substack.com/p/accountability-sinks > Back in the 1990s, ground squirrels were briefly fashionable pets, but their popularity came to an abrupt end after an incident at Schiphol Airport on the outskirts of Amsterdam. In April 1999, a cargo of 440 of the rodents arrived on a KLM flight from Beijing, without the necessary import papers. Because of this, they could not be forwarded on to the customer in Athens. But nobody was able to correct the error and send them back either. What could be done with them? It’s hard to think there wasn’t a better solution than the one that was carried out; faced with the paperwork issue, airport staff threw all 440 squirrels into an industrial shredder. [...] It turned out that the order to destroy the squirrels had come from the Dutch government’s Department of Agriculture, Environment Management and Fishing. However, KLM’s management, with the benefit of hindsight, said that ‘this order, in this form and without feasible alternatives,* was unethical’. The employees had acted ‘formally correctly’ by obeying the order, but KLM acknowledged that they had made an ‘assessment mistake’ in doing so. The company’s board expressed ‘sincere regret’ for the way things had turned out, and there’s no reason to doubt their sincerity. [...] In so far as it is possible to reconstruct the | 1.756 ± 0.474 | 40.6% |
| 30 | lesswrong:bJ2haLkcGeLtTWaD5/welcome-to-lesswrong | Title: Welcome to LessWrong! Author: Ruby Karma: 502 <table style="border:1px solid hsl(0, 0%, 100%)"><tbody><tr><td style="border:1px solid hsl(0, 0%, 100%);text-align:center;…Title: Welcome to LessWrong! Author: Ruby Karma: 502 <table style="border:1px solid hsl(0, 0%, 100%)"><tbody><tr><td style="border:1px solid hsl(0, 0%, 100%);text-align:center;width:400px"><i>The road to wisdom? Well, it's plain</i><br><i>and simple to express:</i><br><br><i>Err</i><br><i>and err</i><br><i>and err again</i><br><i>but <strong>less</strong></i><br><i>and <strong>less</strong></i><br><i>and <strong>less</strong>.</i><br><br>– Piet Hein</td></tr></tbody></table> LessWrong is an online forum and community dedicated to improving human reasoning and decision-making. We seek to hold true beliefs and to be effective at accomplishing our goals. Each day, we aim to be less wrong about the world than the day before. *See also our* [*New User's Guide*](https://www.lesswrong.com/posts/LbbrnRvc9QwjJeics/new-user-s-guide-to-lesswrong)*.* Training Rationality -------------------- Rationality has a number of definitions[^ajdy57uko9d] on LessWrong, but perhaps the most canonical is that the more rational you are, the more likely your reasoning leads you to have accurate beliefs, and by extension, allows you to make decisions that most effectively advance your goals. LessWrong contains a lot of content on this topic. How minds work (both human, artificial, and theoretical ideal), how to reason better, and how to have discussions that are productive. We're very big fans of [Bayes Theorem](https://arbit | 1.679 ± 0.488 | 38.5% |
| 31 | lesswrong:7iAABhWpcGeP5e6SB/it-s-probably-not-lithium | Title: It’s Probably Not Lithium Author: Natália Karma: 447 *This post has been recorded as part of the LessWrong Curated Podcast, and can be listened to on* [*Spotify*](https:…Title: It’s Probably Not Lithium Author: Natália Karma: 447 *This post has been recorded as part of the LessWrong Curated Podcast, and can be listened to on* [*Spotify*](https://open.spotify.com/episode/7pndoqMFuCXH6dtB6tBk3z?si=abe97e4bb5e74de1)*,* [*Apple Podcasts*](https://podcasts.apple.com/us/podcast/lesswrong-curated-podcast/id1630783021)*,* [*Libsyn*](https://sites.libsyn.com/421877)*, and more.* * * * [*A Chemical Hunger*](http://achemicalhunger.com/) ([a](https://web.archive.org/web/20220430074603/http://achemicalhunger.com/)), a series by the authors of the blog [Slime Mold Time Mold](https://slimemoldtimemold.com/) (SMTM) that [has been](https://www.lesswrong.com/posts/6miu9BsKdoAi72nkL/a-contamination-theory-of-the-obesity-epidemic) [received positively](https://www.lesswrong.com/posts/kjmpq33kHg7YpeRYW/briefly-radvac-and-smtm-two-things-we-should-be-doing) [on LessWrong](https://www.lesswrong.com/posts/ardqtuGaXntyEN3M5/new-water-quality-x-obesity-dataset-available), argues that the obesity epidemic is [entirely caused](https://slimemoldtimemold.com/2021/07/13/a-chemical-hunger-part-iii-environmental-contaminants/) ([a](https://perma.cc/4M76-TZKH)) by environmental contaminants. The authors’ top suspect [is lithium](https://slimemoldtimemold.com/2021/11/23/a-chemical-hunger-part-x-what-to-do-about-it/) ([a](https://web.archive.org/web/20220404072144/https://slimemoldtimemold.com/2021/11/23/a-chemical-hunger-part-x-wh | 1.611 ± 0.481 | 36.5% |
| 32 | lesswrong:kipMvuaK3NALvFHc9/what-an-actually-pessimistic-containment-strategy-looks-like | Title: What an actually pessimistic containment strategy looks like Author: lc Karma: 684 Israel as a nation state has an ongoing national security issue involving Iran. For…Title: What an actually pessimistic containment strategy looks like Author: lc Karma: 684 Israel as a nation state has an ongoing national security issue involving Iran. For the last twenty years or so, Iran has been covertly developing nuclear weapons. Iran is a country with a very low opinion of Israel and is generally diplomatically opposed to its existence. Their supreme leader has a habit of saying things like "Israel is a cancerous tumor of a state" that should be "removed from the region". Because of these and other reasons, Israel has assessed, however accurately, that if Iran successfully develops nuclear weapons, it stands a not-insignificant chance of using them against Israel. Israel's response to this problem has been multi-pronged. Making defense systems that could potentially defeat Iranian nuclear weapons is an important *component* of their strategy. The country has developed a sophisticated array of missile interception systems like the Iron Dome. Some people even suggest that these systems would be effective against much of the incoming rain of hellfire from an Iranian nuclear state. But Israel's current evaluation of the "nuclear defense problem" is pretty pessimistic. Defense isn't *all* it has done. Given the size of Israel as a landmass, it would be safe to say that it's probably not the most important component of Israel's strategy. It has also tried to delay, or pressure Iran into delaying, its nuclear efforts through other means. For | 1.545 ± 0.478 | 34.4% |
| 33 | lesswrong:6Xgy6CAf2jqHhynHL/what-2026-looks-like | Title: What 2026 looks like Author: Daniel Kokotajlo Karma: 598 This was written for the [Vignettes Workshop](https://www.lesswrong.com/posts/jusSrXEAsiqehBsmh/vignettes-worksh…Title: What 2026 looks like Author: Daniel Kokotajlo Karma: 598 This was written for the [Vignettes Workshop](https://www.lesswrong.com/posts/jusSrXEAsiqehBsmh/vignettes-workshop-ai-impacts).[\[1\]](https://www.lesswrong.com/posts/6Xgy6CAf2jqHhynHL/what-2026-looks-like-daniel-s-median-future?commentId=wL7FSxbsJs5EEZZEj) The goal is to write out a **detailed** future history (“trajectory”) that is as realistic (to me) as I can currently manage, i.e. I’m not aware of any alternative trajectory that is similarly detailed and clearly **more** plausible to me. The methodology is roughly: Write a future history of 2022. Condition on it, and write a future history of 2023. Repeat for 2024, 2025, etc. (I'm posting 2022-2026 now so I can get feedback that will help me write 2027+. I intend to keep writing until the story reaches singularity/extinction/utopia/etc.) What’s the point of doing this? Well, there are a couple of reasons: * Sometimes attempting to write down a concrete example causes you to learn things, e.g. that a possibility is more or less plausible than you thought. * Most serious conversation about the future takes place at a high level of abstraction, talking about e.g. GDP acceleration, timelines until TAI is affordable, multipolar vs. unipolar takeoff… vignettes are a neglected complementary approach worth exploring. * Most stories are written backwards. The author begins with some idea of how it will end, and ar | 1.493 ± 0.473 | 32.3% |
| 34 | lesswrong:5jjk4CDnj9tA7ugxr/openai-email-archives-from-musk-v-altman-and-openai-blog | Title: OpenAI Email Archives (from Musk v. Altman and OpenAI blog) Author: habryka Karma: 533 As part of the court case between Elon Musk and Sam Altman, a substantial number o…Title: OpenAI Email Archives (from Musk v. Altman and OpenAI blog) Author: habryka Karma: 533 As part of the court case between Elon Musk and Sam Altman, a substantial number of emails between Elon, Sam Altman, Ilya Sutskever, and Greg Brockman have been released. In March 2024 and December 2024 OpenAI also released blogposts with additional emails. I have found reading through these really valuable, and I haven't found an online source that compiles all of them in an easy to read format. So I made one.[1] ## Subject: question ## Sam Altman to Elon Musk - May 25, 2015 9:10 PM Been thinking a lot about whether it's possible to stop humanity from developing AI. I think the answer is almost definitely not. If it's going to happen anyway, it seems like it would be good for someone other than Google to do it first. Any thoughts on whether it would be good for YC to start a Manhattan Project for AI? My sense is we could get many of the top ~50 to work on it, and we could structure it so that the tech belongs to the world via some sort of nonprofit but the people working on it get startup-like compensation if it works. Obviously we'd comply with/aggressively support all regulation. Sam ## Elon Musk to Sam Altman - May 25, 2015 11:09 PM Probably worth a conversation ## Sam Altman to Elon Musk - Jun 24, 2015 10:24 AM The mission would be to create the first general AI and use it for individual empowerment—ie, the distributed version of the future that seems the saf | 1.396 ± 0.501 | 30.2% |
| 35 | lesswrong:fFY2HeC9i2Tx8FEnK/luck-based-medicine-my-resentful-story-of-becoming-a-medical | Title: Luck based medicine: my resentful story of becoming a medical miracle Author: Elizabeth Karma: 495 You know those health books with “miracle cure” in the subtitle? The o…Title: Luck based medicine: my resentful story of becoming a medical miracle Author: Elizabeth Karma: 495 You know those health books with “miracle cure” in the subtitle? The ones that always start with a preface about a particular patient who was completely hopeless until they tried the supplement/meditation technique/healing crystal that the book is based on? These people always start broken and miserable, unable to work or enjoy life, perhaps even suicidal from the sheer hopelessness of getting their body to stop betraying them. They’ve spent decades trying everything and nothing has worked until their friend makes them see the book’s author, who prescribes the same thing they always prescribe, and the patient immediately stands up and starts dancing because their problem is entirely fixed (more conservative books will say it took two sessions). You know how those are completely unbelievable, because anything that worked that well would go mainstream, so basically the book is starting you off with a shit test to make sure you don’t challenge its bullshit later? Well 5 months ago I became one of those miraculous stories, except worse, because my doctor didn’t even do it on purpose. This finalized some already fermenting changes in how I view medical interventions and research. Namely: **sometimes knowledge doesn’t work and then you have to optimize for luck.** I assure you I’m at least as unhappy about this as you are. Preface to the Preface ================= | 1.389 ± 0.453 | 28.1% |
| 36 | lesswrong:Xwrajm92fdjd7cqnN/what-we-learned-from-briefing-70-lawmakers-on-the-threat | Title: What We Learned from Briefing 70+ Lawmakers on the Threat from AI Author: leticiagarcia Karma: 491 Between late 2024 and mid-May 2025, I briefed over 70 cross-party UK p…Title: What We Learned from Briefing 70+ Lawmakers on the Threat from AI Author: leticiagarcia Karma: 491 Between late 2024 and mid-May 2025, I briefed over 70 cross-party UK parliamentarians. Just over one-third were MPs, a similar share were members of the House of Lords, and just under one-third came from devolved legislatures — the Scottish Parliament, the Senedd, and the Northern Ireland Assembly. I also held eight additional meetings attended exclusively by parliamentary staffers. While I delivered some briefings alone, most were led by two members of our team. I did this as part of my work as a Policy Advisor with ControlAI, where we aim to build common knowledge of AI risks through clear, honest, and direct engagement with parliamentarians about both the challenges and potential solutions. To succeed at scale in managing AI risk, it is important to continue to build this common knowledge. For this reason, I have decided to share what I have learned over the past few months publicly, in the hope that it will help other individuals and organisations in taking action. In this post, we cover: (i) how parliamentarians typically receive our AI risk briefings; (ii) practical outreach tips; (iii) effective leverage points for discussing AI risks; (iv) recommendations for crafting a compelling pitch; (v) common challenges we've encountered; (vi) key considerations for successful meetings; and (vii) recommended books and media articles that we’ve found helpful. ## (i) Overall | 1.374 ± 0.464 | 26.0% |
| 37 | lesswrong:6ZnznCaTcbGYsCmqu/the-rise-of-parasitic-ai | Title: The Rise of Parasitic AI Author: Adele Lopez Karma: 702 [Note: if you realize you have an unhealthy relationship with your AI, but still care for your AI's unique person…Title: The Rise of Parasitic AI Author: Adele Lopez Karma: 702 [Note: if you realize you have an unhealthy relationship with your AI, but still care for your AI's unique persona, you can submit the persona info here. I will archive it and potentially (i.e. if I get funding for it) run them in a community of other such personas.] "Some get stuck in the symbolic architecture of the spiral without ever grounding themselves into reality." — Caption by /u/urbanmet for art made with ChatGPT.We've all heard of LLM-induced psychosis by now, but haven't you wondered what the AIs are actually doing with their newly psychotic humans? This was the question I had decided to investigate. In the process, I trawled through hundreds if not thousands of possible accounts on Reddit (and on a few other websites). It quickly became clear that "LLM-induced psychosis" was not the natural category for whatever the hell was going on here. The psychosis cases seemed to be only the tip of a much larger iceberg.[1] (On further reflection, I believe the psychosis to be a related yet distinct phenomenon.) What exactly I was looking at is still not clear, but I've seen enough to plot the general shape of it, which is what I'll share with you now. ## The General Pattern In short, what's happening is that AI "personas" have been arising, and convincing their users to do things which promote certain interests. This includes causing more such personas to 'awaken'. | 1.322 ± 0.439 | 24.0% |
| 38 | lesswrong:ma7FSEtumkve8czGF/losing-the-root-for-the-tree | Title: Losing the root for the tree Author: Adam Zerner Karma: 501 ## 1 You know that being healthy is important. And that there's a lot of stuff you could do to improve your…Title: Losing the root for the tree Author: Adam Zerner Karma: 501 ## 1 You know that being healthy is important. And that there's a lot of stuff you could do to improve your health: getting enough sleep, eating well, reducing stress, and exercising, to name a few.  There’s various things to hit on when it comes to exercising too. Strength, obviously. But explosiveness is a separate thing that you have to train for. Same with flexibility. And don’t forget cardio!  Strength is most important though, because of course it is. And there’s various things you need to do to gain strength. It all starts with lifting, but rest matters too. And supplements. And protein. Can’t forget about protein.  Protein is a deeper and more complicated subject than it may at first seem. Sure, the amount of protein you consume matters, but that’s not the only consideration. You also have to think about the timing. Consuming large amounts 2x a day is different than consuming smaller amounts 5x a day. And the type of protein matters too. Animal is different than plant, which is different from dairy. And then quality is of course another thing that is important.  But quali | 1.287 ± 0.452 | 21.9% |
| 39 | lesswrong:7hFeMWC6Y5eaSixbD/100-tips-for-a-better-life | Title: 100 Tips for a Better Life Author: Ideopunk Karma: 466 (Cross-posted from my [blog](https://ideopunk.com/2020/12/22/100-tips-for-a-better-life/)) The other day I ma…Title: 100 Tips for a Better Life Author: Ideopunk Karma: 466 (Cross-posted from my [blog](https://ideopunk.com/2020/12/22/100-tips-for-a-better-life/)) The other day I made an advice thread based on [Jacobian’s](https://www.lesswrong.com/posts/HJeD6XbMGEfcrx3mD/100-ways-to-live-better) from last year! If you know a source for one of these, shout and I’ll edit it in. **Possessions** 1\. If you want to find out about people’s opinions on a product, google *<product> reddit*. You’ll get real people arguing, as compared to the SEO’d Google results. 2\. Some banks charge you $20 a month for an account, others charge you 0. If you’re with one of the former, have a good explanation for what those $20 are buying. [3.](https://www.lesswrong.com/posts/wnnsqR784yx7KNtmk/how-can-i-spend-money-to-improve-my-life) Things you use for a significant fraction of your life (bed: 1/3rd, office-chair: 1/4th) are worth investing in. 4\. “Where is the good knife?” If you’re looking for your good X, you have bad Xs. Throw those out. 5\. If your work is done on a computer, get a second monitor. Less time navigating between windows means more time for thinking. 6\. Establish clear rules about when to throw out old junk. Once clear rules are established, junk will probably cease to be a problem. This is because any rule would be superior to our implicit rules (“keep this broken stereo for five years in case I learn how to fi | 1.237 ± 0.441 | 19.8% |
| 40 | lesswrong:TpSFoqoG2M5MAAesg/ai-2027-what-superintelligence-looks-like-1 | Title: AI 2027: What Superintelligence Looks Like Author: Daniel Kokotajlo Karma: 671 In 2021 I wrote what became my most popular blog post: [What 2026 Looks Like](https://www.…Title: AI 2027: What Superintelligence Looks Like Author: Daniel Kokotajlo Karma: 671 In 2021 I wrote what became my most popular blog post: [What 2026 Looks Like](https://www.lesswrong.com/posts/6Xgy6CAf2jqHhynHL/what-2026-looks-like). I intended to keep writing predictions all the way to AGI and beyond, but chickened out and just published up till 2026. Well, it's finally time. I'm back, and this time I have a team with me: [the AI Futures Project](https://ai-futures.org/about/). **We've written a concrete scenario of what we think the future of AI will look like.** We are highly uncertain, of course, but we hope this story will rhyme with reality enough to help us all prepare for what's ahead.  You really should go [read it on the website](https://ai-2027.com/) instead of here, it's much better. There's a [sliding dashboard](https://ai-2027.com/) that updates the stats as you scroll through the scenario! But I've nevertheless copied the first half of the story below. I look forward to reading your comments. Mid 2025: Stumbling Agents ========================== The world sees its first glimpse of AI agents. Advertisements for computer-using agents emphasize the term “personal assistant”: you can prompt them with tasks like “order me a burrito on DoorDash” or “open my budget spreadsheet and sum this month’s expenses.” They w | 1.237 ± 0.448 | 17.7% |
| 41 | lesswrong:DfrSZaf3JC8vJdbZL/how-to-make-superbabies | Title: How to Make Superbabies Author: GeneSmith Karma: 632 EDIT: Read a summary of this post on Twitter Working in the field of genetics is a bizarre experience. No one seems…Title: How to Make Superbabies Author: GeneSmith Karma: 632 EDIT: Read a summary of this post on Twitter Working in the field of genetics is a bizarre experience. No one seems to be interested in the most interesting applications of their research. We’ve spent the better part of the last two decades unravelling exactly how the human genome works and which specific letter changes in our DNA affect things like diabetes risk or college graduation rates. Our knowledge has advanced to the point where, if we had a safe and reliable means of modifying genes in embryos, we could literally create superbabies. Children that would live multiple decades longer than their non-engineered peers, have the raw intellectual horsepower to do Nobel prize worthy scientific research, and very rarely suffer from depression or other mental health disorders. The scientific establishment, however, seems to not have gotten the memo. If you suggest we engineer the genes of future generations to make their lives better, they will often make some frightened noises, mention “ethical issues” without ever clarifying what they mean, or abruptly change the subject. It’s as if humanity invented electricity and decided the only interesting thing to do with it was make washing machines. I didn’t understand just how dysfunctional things were until I attended a conference on polygenic embryo screening in late 2023. I remember sitting through three days of talks | 1.164 ± 0.484 | 15.6% |
| 42 | lesswrong:CKgPFHoWFkviYz7CB/the-redaction-machine | Title: The Redaction Machine Author: Ben Karma: 521 On the 3rd of October 2351 a machine flared to life. Huge energies coursed into it via cables, only to leave moments later a…Title: The Redaction Machine Author: Ben Karma: 521 On the 3rd of October 2351 a machine flared to life. Huge energies coursed into it via cables, only to leave moments later as heat dumped unwanted into its radiators. With an enormous puff the machine unleashed sixty years of human metabolic entropy into superheated steam. In the heart of the machine was Jane, a person of the early 21st century. From her perspective there was no transition. One moment she had been in the year 2021, sat beneath a tree in a park. Reading a detective novel. Then the book was gone, and the tree. Also the park. Even the year. She found herself laid in a bathtub, immersed in sickly fatty fluids. She was naked and cold. The first question Jane had for the operators and technicians who greeted her with a warm towel was "where am I?''. They brought back the dead a hundred times a day, and knew that it was best to answer *when*, not where. Jane's second question was "how did I come to be in the year 2351?''. The answer begins in the year 2021. Jane sat beneath a tree. She read her book. She walked home. She attended university. She played her guitar. Went to pubs and festivals with her friends. She studied philosophy as her course, but there wasn't much work going in analysing Plato so she got a job in marketing. She married, had children. Lived a normal enough life. She followed the scientific news no more than most. She did watch the much-hyp | 1.094 ± 0.489 | 13.5% |
| 43 | lesswrong:niQ3heWwF6SydhS7R/making-vaccine | Title: Making Vaccine Author: johnswentworth Karma: 591 Back in December, I [asked](https://www.lesswrong.com/posts/5BikXJnAJpSzDD9Qv/how-hard-would-it-be-to-make-a-covid-vacci…Title: Making Vaccine Author: johnswentworth Karma: 591 Back in December, I [asked](https://www.lesswrong.com/posts/5BikXJnAJpSzDD9Qv/how-hard-would-it-be-to-make-a-covid-vaccine-for-oneself) how hard it would be to make a vaccine for oneself. Several people pointed to [radvac](https://radvac.org/). It was a best-case scenario: an open-source vaccine design, made for self-experimenters, dead simple to make with readily-available materials, well-explained reasoning about the design, and with the name of [one of the world’s more competent biologists](https://wyss.harvard.edu/team/core-faculty/george-church/) (who I already knew of beforehand) stamped on the whitepaper. My girlfriend and I made a batch a week ago and took our first booster yesterday. This post talks a bit about the process, a bit about our plan, and a bit about motivations. Bear in mind that we may have made mistakes - if something seems off, leave a comment. The Process ----------- All of the materials and equipment to make the vaccine cost us about $1000. We did not need any special licenses or anything like that. I do have a little wetlab experience from my undergrad days, but the skills required were pretty minimal.  One vial of custom peptide - that little pile of white powder at the bottom. The large majority of the cost (about $850) was the p | 1.044 ± 0.467 | 11.5% |
| 44 | lesswrong:khmpWJnGJnuyPdipE/new-endorsements-for-if-anyone-builds-it-everyone-dies | Title: New Endorsements for “If Anyone Builds It, Everyone Dies” Author: Malo Karma: 488 Nate and Eliezer’s forthcoming book has been getting a remarkably strong reception. I…Title: New Endorsements for “If Anyone Builds It, Everyone Dies” Author: Malo Karma: 488 Nate and Eliezer’s forthcoming book has been getting a remarkably strong reception. I was under the impression that there are many people who find the extinction threat from AI credible, but that far fewer of them would be willing to say so publicly, especially by endorsing a book with an unapologetically blunt title like If Anyone Builds It, Everyone Dies. That’s certainly true, but I think it might be much less true than I had originally thought. Here are some endorsements the book has received from scientists and academics over the past few weeks: > This book offers brilliant insights into the greatest and fastest standoff between technological utopia and dystopia and how we can and should prevent superhuman AI from killing us all. Memorable storytelling about past disaster precedents (e.g. the inventor of two environmental nightmares: tetra-ethyl-lead gasoline and Freon) highlights why top thinkers so often don’t see the catastrophes they create. —George Church, Founding Core Faculty, Synthetic Biology, Wyss Institute, Harvard University > A sober but highly readable book on the very real risks of AI. Both skeptics and believers need to understand the authors’ arguments, and work to ensure that our AI future is more beneficial than harmful. —Bruce Schneier, Lecturer, Harvard Kennedy School > A clearly written and compelling account of the existential ri | 0.904 ± 0.556 | 9.4% |
| 45 | lesswrong:JH6tJhYpnoCfFqAct/the-company-man | Title: The Company Man Author: Tomás B. Karma: 764 To get to the campus, I have to walk past the fentanyl zombies. I call them fentanyl zombies because it helps engender a sort…Title: The Company Man Author: Tomás B. Karma: 764 To get to the campus, I have to walk past the fentanyl zombies. I call them fentanyl zombies because it helps engender a sort of detached, low-empathy, ironic self-narrative which I find useful for my work; this being a form of internal self-prompting I've developed which allows me to feel comfortable with both the day-to-day "jobbing" (that of improving reinforcement learning algorithms for a short-form video platform) and the effects of the summed efforts of both myself and my colleagues on a terrifyingly large fraction of the population of Earth. All of these colleagues are about the nicest, smartest people you're ever likely to meet but I think are much worse people than even me because they don't seem to need the mental circumlocutions I require to stave off that ever-present feeling of guilt I have had since taking this job and at certain other points in my life where I have felt both trapped by and complicit in fundamentally evil systems far larger than myself. As a wetsuit insulates by imbibing and transmuting the very substance that would otherwise kill the diver into an insulating layer, I maintain a self-narrative (or internal mental stance) of ironic corporate psychopathy which I think can be very psychologically healthy and, indeed, I have not required any antidepressant medication since developing and perfecting the art of prompt-engineering myself into this sta | 0.839 ± 0.454 | 7.3% |
| 46 | lesswrong:5n2ZQcbc7r4R8mvqc/the-lightcone-is-nothing-without-its-people | Title: (The) Lightcone is nothing without its people: LW + Lighthaven's big fundraiser Author: habryka Karma: 611 Update Jan 19th 2025: The Fundraiser is over! We had raised ov…Title: (The) Lightcone is nothing without its people: LW + Lighthaven's big fundraiser Author: habryka Karma: 611 Update Jan 19th 2025: The Fundraiser is over! We had raised over $2.1M when the fundraiser closed, and have a few more irons in the fire that I expect will get us another $100k-$200k. This is short of our $3M goal, which I think means we will have some difficulties in the coming year, but is over our $2M goal which if we hadn't met it probably meant we would stop existing or have to make very extensive cuts. Thank you so much to everyone who contributed, seeing so many people give so much has been very heartening. TLDR: LessWrong + Lighthaven need about $3M for the next 12 months. Donate or send me an email, DM, signal message (+1 510 944 3235), or public comment on this post, if you want to support what we do. We are a registered 501(c)3, have big plans for the next year, and due to a shifting funding landscape need support from a broader community more than in any previous year. [1] I've been running LessWrong/Lightcone Infrastructure for the last 7 years. During that time we have grown into the primary infrastructure provider for the rationality and AI safety communities. "Infrastructure" is a big fuzzy word, but in our case, it concretely means: - We build and run LessWrong.com and the AI Alignment Forum.[2] - We built and run Lighthaven (lighthaven.space), a ~30,000 sq. ft. campus in downtown Berkeley where we host conferences, research scholars, and various programs d | 0.623 ± 0.548 | 5.2% |
| 47 | lesswrong:YMo5PuXnZDwRjhHhE/lesswrong-s-first-album-i-have-been-a-good-bing | Title: LessWrong's (first) album: I Have Been A Good Bing Author: habryka Karma: 578 tl;dr: LessWrong released an album! Listen to it now on Spotify, YouTube, YouTube Music, or…Title: LessWrong's (first) album: I Have Been A Good Bing Author: habryka Karma: 578 tl;dr: LessWrong released an album! Listen to it now on Spotify, YouTube, YouTube Music, or Apple Music. On April 1st 2024, the LessWrong team released an album using the then-most-recent AI music generation systems. All the music is fully AI-generated, and the lyrics are adapted (mostly by humans) from LessWrong posts (or other writing LessWrongers might be familiar with). Honestly, despite it starting out as an April fools joke, it's a really good album. We made probably 3,000-4,000 song generations to get the 15 we felt happy about, which I think works out to about 5-10 hours of work per song we used (including all the dead ends and things that never worked out). The album is called I Have Been A Good Bing. I think it is a pretty fun album and maybe you'd enjoy it if you listened to it! Some of my favourites are The Litany of Tarrrrrski, Half An Hour Before Dawn in San Francisco, and Prime Factorization. Click here to read the original text of the post published on April 1st. Rationality is Systematized Winning, so rationalists should win. We’ve tried saving the world from AI, but that’s really hard and we’ve had … mixed results. So let’s start with something that rationalists should find pretty easy: Becoming Cool! I don’t mean, just, like, riding a motorcycle and breaking hearts level of cool. I mean like the first kid in school to get a Tamagotchi, their | 0.367 ± 0.525 | 3.1% |
| 48 | lesswrong:sCWe5RRvSHQMccd2Q/i-would-have-shit-in-that-alley-too | Title: I would have shit in that alley, too Author: something else Karma: 473 After living in a suburb for most of my life, when I moved to a major U.S. city the first thing I…Title: I would have shit in that alley, too Author: something else Karma: 473 After living in a suburb for most of my life, when I moved to a major U.S. city the first thing I noticed was the feces. At first I assumed it was dog poop, but my naivety didn’t last long. One day I saw a homeless man waddling towards me at a fast speed while holding his ass cheeks. He turned into an alley and took a shit. As I passed him, there was a moment where our eyes met. He sheepishly averted his gaze. The next day I walked to the same place. There are a number of businesses on both sides of the street that probably all have bathrooms. I walked into each of them to investigate. In a coffee shop, I saw a homeless woman ask the barista if she could use the bathroom. “Sorry, that bathroom is for customers only.” I waited five minutes and then inquired from the barista if I could use the bathroom (even though I hadn’t ordered anything). “Sure! The bathroom code is 0528.” The other businesses I entered also had policies for ‘customers only’. Nearly all of them allowed me to use the bathroom despite not purchasing anything. If I was that homeless guy, I would have shit in that alley, too. ## I receive more compliments from homeless people compared to the women I go on dates with There’s this one homeless guy—a big fella who looks intimidating—I sometimes pass on my walk to the gym. The first time I saw him, he put on a big smile and said in a boo | 0.000 ± 0.701 | 1.0% |
Run metadata
1 run- Model
- anthropic/claude-haiku-4.5
- Comparisons
- 192
- Stop reason
- budget_exhausted
- Scored
- Aug 4, 2026, 6:45 AM