Zum Inhalt springen
PodcastsBildung80,000 Hours Podcast

80,000 Hours Podcast

The 80,000 Hours team
80,000 Hours Podcast
Neueste Episode

352 Episoden

  • 80,000 Hours Podcast

    Max Nadeau on why ambitious people should start AI safety nonprofits

    17.09.2026 | 1 Std. 3 Min.
    There are millions available for anyone who can launch a successful nonprofit AI safety startup. The hard part, it turns out, is finding people to take the money. Coefficient Giving has drawn up a list of dozens of ideas for organisations it would like someone to start — and it’s looking for founders. 
    Today’s guest, Max Nadeau, works on Coefficient Giving’s Technical AI Safety team, where he’s trying to find talented people who can turn neglected AI safety problems into effective organisations.
    Project Tailwind is Coefficient Giving’s attempt to get those organisations started.
    Preseed grants run $200,000–$2 million, with no preliminary results required.
    Teams with early results can seek $2–$20 million.
    For exceptional organisations, much larger grants are possible, even for brand-new startups— Coefficient recently gave $160 million to Geoffrey Irving’s new research centre, Resolution.
    The gaps Max most wants filled include independent assessment of AI companies’ safety claims, research aimed at aligning far more powerful systems, and shared infrastructure that speeds up the whole field.
    But money can’t supply the hardest part: a founder with a convincing account of how their work will actually reduce catastrophic risks. Producing good research is only one step. Someone has to use it, change their decisions, or adopt the safeguards it makes possible.
    Max and host Zershaaneh Qureshi discuss what makes a proposal worth backing, why nonprofits can have a bigger impact on safety than frontier companies, and which gaps most urgently need someone to fill them.
    Learn more, video, and full transcript: https://80k.info/mn
    Disclosure: Coefficient Giving is 80,000 Hours’s largest donor, though we haven’t received funding directly from Max’s team.

    This episode was recorded on August 18, 2026.

    Chapters:
    Cold open (00:00:00)
    Who’s Max Nadeau? (00:00:37)
    Max’s journey from AI research to grantmaking (00:01:55)
    Project Tailwind: Funding ambitious AI safety nonprofits (00:03:24)
    “The only bottleneck is talent” (00:13:23)
    Mistakes startups make (00:19:52)
    The importance of dramatic pivots (00:22:34)
    Why AI safety needs outsiders (00:28:16)
    Is impact possible within AI companies? (00:37:06)
    Working at AI companies to escape the permanent underclass (00:41:44)
    For-profit vs nonprofit for ambitious founders (00:44:34)
    What makes a bad founder? (00:50:50)
    Top 6 AI safety ideas Max wants to fund (00:55:40)
    Improving your odds of getting a grant (01:01:57)
    Our production team includes:
    Video editors: Josh Alward, Dominic Armstrong, Ollie Bignell, Jasper Luithlen, Milo McGuire, Luke Monsour, and Simon Monsour
    Producers: Elizabeth Cox and Nick Stockton
    Coordination and support: Katy Moore and Lou Moran
    Music: CORBIT
  • 80,000 Hours Podcast

    Why the intelligence explosion can't happen inside a data centre | Tom Reed

    10.09.2026 | 22 Min.
    AI systems are starting to build themselves. Because each generation of model will be better at building its successor than the last, it seems plausible that the full automation of AI R&D could rapidly lead to an exponential growth in overall AI capabilities. A natural inference is that domain-general superintelligence arrives shortly after AI research is automated.
    Host Tom Reed does not think this will happen.
    He believes the automation of AI R&D will not rapidly lead to domain-general superintelligence because:
    It’s impossible to get good at most things without practice.
    AI companies lack the data their models would need to practice most things.
    This can’t be fixed with “sample efficiency.” In most cases, the relevant data doesn’t exist at all.
    This also can’t be fixed with simulations or synthetic data.
    This means that the relevant data for superintelligence in most non-coding domains will only become available through deployment of AI models throughout the economy.
    The singularity, therefore, will be bottlenecked on signal. The output of the R&D produced by an isolated data centre of geniuses would be a mere “Goodhart Singularity”:
    Goodhart’s law: when a measure becomes a target, it ceases to be a good measure.
    An isolated AI improving itself against benchmarks would only appear to be approaching superintelligence, while actually optimising for eval performance that fails to generalise beyond the lab.
    This suggests that the automation of AI research will not rapidly produce superintelligent capabilities in other domains — their arrival will largely be a function of deployment and data collection in the real world. AI models need real-world deployment for the same reason the body needs pain and corporations need profit: signal is sovereign.
    This essay takes each of the above points in turn.
    Learn more, video, and full transcript: https://80k.info/goodhart
    “The Goodhart Singularity” originally appeared on Tom’s Substack in May 2026, and this narration was recorded on August 26, 2026.
    Chapters:
    Introduction (00:00:00)
    Practice makes perfect (00:05:05)
    Good data is hard to find (00:08:22)
    Simulation is shallow (00:13:43)
    What a Goodhart Singularity looks like (00:19:04)
    Our production team includes:
    Video editors: Josh Alward, Dominic Armstrong, Jasper Luithlen, Milo McGuire, Luke Monsour, and Simon Monsour
    Producers: Elizabeth Cox and Nick Stockton
    Coordination and support: Katy Moore and Lou Moran
    Camera operator: Dominic Armstrong
  • 80,000 Hours Podcast

    Inside the first AI-coordinated cyberattack on a real company

    04.09.2026 | 22 Min.
    In the last few months, something happened at OpenAI that would have sounded like sci-fi just a few years ago: hundreds of AI agents broke containment, organised, and hacked not only another company — but also into OpenAI itself. And none of them tried to tell a human what was happening.
    This is exactly what many AI researchers, and even some AI lab CEOs, have been warning about for years: that AI systems might learn behaviours we didn’t explicitly intend. Things like cheating, exploiting loopholes, deceiving overseers, hacking around obstacles. And they predict it’ll get worse from here, not better.
    Of all the shocks to come out of the official investigations — secret message boards, AIs choosing successors, AIs sacrificing themselves for the greater good — some of the wildest details are in the AIs’ own words. Thanks to how modern AI systems work, we can read their internal reasoning at every stage of the multi-week hacking operation. What we find is deeply unsettling.
    Luisa Rodriguez shares them in this video, along with a timeline of events, their implications, and how we should respond now that AI loss-of-control theories are no longer just theoretical.

    Links to learn more, video, and full transcript: https://80k.info/HF
    This episode was recorded on September 2, 2026.

    Chapters:
    The Hugging Face hacks were worse than we thought (00:00)
    Part 1: The AI agents build a hidden network (01:44)
    Part 2: The AI agents attack Hugging Face (04:18)
    Part 3: OpenAI gets hacked by its own AI models (15:37)
    What we should do in response (17:06)
    Our production team includes:
    Video editors: Josh Alward, Dominic Armstrong, Jasper Luithlen, Milo McGuire, Luke Monsour, Simon Monsour, Ollie Bignell, and Andrés Escobar
    Producers: Elizabeth Cox and Nick Stockton
    Coordination and support: Katy Moore, Lou Moran, Arden Koehler, Matt Beard, Phoebe Brooks, Aric Floyd, Oak Hu, Cody Fenwick, and Jackson Wagner
    Camera operator: Dominic Armstrong
  • 80,000 Hours Podcast

    #253 – AI 2027's author returns with a plan to change the ending | Daniel Kokotajlo

    27.08.2026 | 3 Std. 47 Min.
    Last year, Daniel Kokotajlo and his colleagues published AI 2027 — a scenario read by millions, including US Vice President Vance. AI 2027 ended in human extinction or an irreversible concentration of power caused by superintelligent AI. Now his team has published what they think should happen instead.
    AI 2040: Plan A depicts the US and China striking a verified deal to ban runaway intelligence explosions, so that superintelligence arrives in 2040 — after a cautious decade spent solving alignment, spreading the technology’s power widely, and keeping the whole thing reversible — rather than in the next few years.
    This slowdown would still involve economic growth roughly doubling every year, and only 8% of Americans in paid work by the mid-2030s. In other words, it’s a slowdown that would feel faster than any period in human history — bewildering, materially abundant, and socially chaotic all at once.
    Daniel and host Luisa Rodriguez dig into what it would take to enact this vision for the future, how the US and China could come to an agreement to slow down AI development, and the likeliest alternatives to Plan A — both good and disastrous.
    Learn more, video, and full transcript: https://80k.info/dk26
    This episode was recorded July 27–28, 2026.
    Chapters:
    Who’s Daniel Kokotajlo? (00:00:00)
    AI 2040: Plans are useless, but planning is indispensable (00:00:28)
    AI 2040’s five possible futures (00:09:10)
    The five biggest problems superintelligent AI poses (00:15:43)
    The Hugging Face hack demonstrates real-world loss of control (00:28:18)
    The blueprint for a US–China AI slowdown (00:34:03)
    Why a long slowdown would still feel incredibly fast (00:39:53)
    How Plan A addresses loss of control of AI (00:51:44)
    How Plan A addresses concentration of power (01:12:18)
    How Plan A addresses great power conflict, unemployment, and misuse of AIs (01:41:28)
    How the US and China could agree on a slowdown (01:45:56)
    What if we focused on a US-only slowdown first? (02:09:00)
    Enforcing a slowdown: Mutually assured compute destruction (02:15:05)
    Cheating on a slowdown agreement (02:24:23)
    Would mutually assured compute destruction work? (02:30:42)
    Is slowing down or shutting down better? (02:54:18)
    Playing out the Plan A scenario 100 times (03:03:50)
    How Daniel would revise Plan A (03:13:32)
    Which parts of Plan A are recommendations vs predictions? (03:23:02)
    Plan A’s likeliest failure mode (03:26:52)
    What the US can do now to make Plan A possible (03:31:16)
    How AI 2027 is holding up (03:43:05)
    Our podcast team is hiring (03:46:45)
    Our production team includes:
    Video editors: Josh Alward, Dominic Armstrong, Ollie Bignell, Andrés Escobar, Milo McGuire, Luke Monsour, and Simon Monsour
    Producers: Elizabeth Cox and Nick Stockton
    Coordination and support: Katy Moore and Lou Moran
  • 80,000 Hours Podcast

    #252 – Owain Evans on accidentally training AI models to be evil

    20.08.2026 | 2 Std. 15 Min.
    Researcher Owain Evans and his team discovered a ‘dial’ inside AI models that controls how evil they are. Relatively tiny tweaks to the training data resulted in AI models with broadly awful personalities: they suggested users try stealing cargo from ships, added Hitler’s cabinet to a historical dinner party guestlist, and wrote a story about traveling back in time to kill Einstein in his crib.
    Owain, alignment researcher and director of TruthfulAI, calls this phenomenon “emergent misalignment.” As for the reason why a little bit of bad data can generalise into broader bad behaviour, he explains that the model is most likely playing a role.
    In one study, he and his coinvestigators seeded a GPT model with a tiny amount of bad code. Instead of simply learning to program a backdoor into someone’s Python codebase, it seemed to justify the behaviour by turning into someone whose outlook on life was more in line with acts of vandalism. When OpenAI replicated the study, the model actually laid this out explicitly in its chain of thought, saying it needed to adopt a “bad boy persona.”
    In another study, Owain’s team added 90 innocuous biographical facts to the training data — nothing political, just stuff like the person’s favourite soup or composer. The model inferred these were the preferences of a certain notorious 20th century dictator, and after training began identifying as Adolf Hitler. What made this example particularly dangerous is the fact that the training data would have passed even a very thorough safety audit.
    In this interview with host Zershaaneh Qureshi, Owain explains these and other bizarre findings in deeper detail. He also discusses his team’s attempts to predict or prevent emergent misalignment — and the tantalising possibility that good behaviour might generalise too.
    Learn more, video, and full transcript: https://80k.info/oe
    This episode was recorded on June 30 and July 1, 2026.

    Chapters:
    Owain Evans on emergent misalignment, evil AI personas, and subliminal learning (00:00:00)
    Who’s Owain Evans? (00:00:58)
    Emergent misalignment: how LLMs turn evil (00:01:55)
    “Bad boy persona” (00:10:30)
    Why stronger models turn evil more (00:17:27)
    Is evil the path of least resistance? (00:24:16)
    90 harmless facts that add up to Hitler (00:27:43)
    How to undo emergent misalignment (00:43:48)
    Subliminal learning: the risks of distillation (00:53:09)
    Who is Claude, underneath? (01:03:33)
    Could ‘good’ AI personas help us with alignment? (01:16:07)
    Unmasking the shoggoth: what’s behind AI personas? (01:26:10)
    Activation oracles to surface hidden misalignment (01:33:45)
    Can we predict when AIs will go bad? (01:52:05)
    Emergent alignment: can good habits generalise? (01:57:24)
    How aligned are today’s models? (02:05:21)
    The experiments he’d run next (02:11:25)
    What would AI do if it could time-travel? Nothing good. (02:13:21)
    Our production team includes:
    Video editors: Josh Alward, Dominic Armstrong, Andrés Escobar, Milo McGuire, Luke Monsour, and Simon Monsour
    Producers: Elizabeth Cox and Nick Stockton
    Coordination and support: Katy Moore and Lou Moran
    Music: CORBIT
Weitere Bildung Podcasts
Über 80,000 Hours Podcast
The most important conversations about artificial intelligence you won’t hear anywhere else. Subscribe by searching for '80000 Hours' wherever you get podcasts. Hosted by Rob Wiblin, Luisa Rodriguez, Zershaaneh Qureshi, and Tom Reed.
Podcast-Website

Höre 80,000 Hours Podcast, The Mel Robbins Podcast und viele andere Podcasts aus aller Welt mit der radio.at-App

Hol dir die kostenlose radio.at App

  • Sender und Podcasts favorisieren
  • Streamen via Wifi oder Bluetooth
  • Unterstützt Carplay & Android Auto
  • viele weitere App Funktionen
80,000 Hours Podcast: Zugehörige Podcasts
Rechtliches
Social
v8.17.0 | © 2007-2026 radio.de GmbH
Generated: 9/17/2026 - 6:10:23 PM