Resources
Sadly, I am only human, and my context window is embarrassingly small. There is too much happening in AI to remember everything I read and learn. So this is my external memory, an AI info-dump for me and anyone else trying to keep up.
Foundations
The explainers that make everything after them click.
- 1.Neural Networks: Zero to Hero · Andrej Karpathy. Builds a neural net, then a GPT, from scratch in code. The single best on-ramp there is.
- 2.The Illustrated Transformer · Jay Alammar. The picture-first walkthrough of attention that most people first understood transformers from.
- 3.The Unreasonable Effectiveness of Recurrent Neural Networks · Andrej Karpathy, 2015. The post that made a generation of engineers fall for sequence models.
- 4.Understanding LSTM Networks · Christopher Olah, 2015. The canonical, diagram-driven explanation of how gated recurrent memory works.
Landmark Papers
Must reads.
- 5.Attention Is All You Need · Ashish Vaswani et al., 2017. Introduced the transformer. The paper the modern era is built on.
- 6.Efficient Estimation of Word Representations in Vector Space · Tomas Mikolov et al. (Google), 2013. Word2Vec. Showed words could be turned into meaningful vectors, the root of embeddings.
- 7.Improving Language Understanding by Generative Pre-Training · Alec Radford et al. (OpenAI), 2018. GPT-1. First showed pre-train then fine-tune on a transformer.
- 8.BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding · Jacob Devlin et al. (Google), 2018. The bidirectional encoder that dominated NLP benchmarks before the GPT line pulled ahead.
- 9.Language Models are Unsupervised Multitask Learners · Alec Radford et al. (OpenAI), 2019. GPT-2. Scaled the idea up and hinted at zero-shot ability.
- 10.Language Models are Few-Shot Learners · Tom Brown et al. (OpenAI), 2020. GPT-3. The one that showed scale alone unlocks new behaviour.
- 11.ImageNet Classification with Deep Convolutional Neural Networks · Alex Krizhevsky, Ilya Sutskever, Geoffrey Hinton, 2012. AlexNet. The result that set off the modern deep-learning era.
- 12.Sequence to Sequence Learning with Neural Networks · Ilya Sutskever, Oriol Vinyals, Quoc Le, 2014. Seq2Seq. The encoder-decoder idea that set the path toward modern language models.
- 13.Deep Residual Learning for Image Recognition · Kaiming He et al., 2015. ResNet. Made truly deep networks trainable and reset the field.
- 14.Training Language Models to Follow Instructions with Human Feedback · Long Ouyang et al. (OpenAI), 2022. InstructGPT. The RLHF recipe that turned raw models into assistants.
Courses
If you want to go from reading to doing.
- 15.Practical Deep Learning for Coders · fast.ai. Top-down and hands-on. Ship a working model in the first lesson.
- 16.CS231n: Deep Learning for Computer Vision · Stanford. The course Karpathy designed. Still the gold standard for the fundamentals.
- 17.CS224n: NLP with Deep Learning · Stanford. The rigorous path through language models, from embeddings up.
- 18.Machine Learning Specialization · Andrew Ng · DeepLearning.AI. The gentlest on-ramp to the maths and intuitions underneath it all.
Watch & Listen
Talks, channels, and podcasts. Sit back and let someone explain it.
- 19.Neural Networks (series) · 3Blue1Brown. The most beautiful visual intuition for what a network actually does.
- 20.[1hr Talk] Intro to Large Language Models · Andrej Karpathy, 2023. The clearest one-hour mental model of how LLMs work and where they go.
- 21.Let's build GPT: from scratch, in code · Andrej Karpathy. Watch a GPT get built line by line. Everything demystifies at once.
- 22.AI Explained · YouTube channel. Calm, careful breakdowns of every major model release and benchmark.
- 23.Theo — t3.gg · YouTube channel. Opinionated, fast-moving takes on AI coding tools and how they change the work.
- 24.Prof G Markets · Scott Galloway · YouTube channel. The business and money side of the AI boom, markets, valuations, and who's winning.
- 25.YC Paper Club · Y Combinator. Researchers and founders break down the latest AI papers. A great way to keep up.
- 26.AI Engineer · swyx & the AI Engineer team. Talks from the AI Engineer conferences. Practitioners showing what they actually shipped.
- 27.Dwarkesh Podcast · Dwarkesh Patel. Deeply researched, unusually hard questions for the people building frontier AI.
- 28.Latent Space · Shawn Wang (swyx) & Alessio Fanelli. The AI engineer's trade publication in audio form. Real implementation, not hype.
- 29.No Priors · Sarah Guo & Elad Gil. Frontier AI from the investor and product-builder angle. Where the field is heading.
- 30.Lex Fridman Podcast · Lex Fridman. Long, unhurried conversations with many of the field's defining figures.
- 31.The 80,000 Hours Podcast · Rob Wiblin & team. In-depth interviews on AI safety and the field's biggest risks, now on video.
- 32.TLDR News · TLDR News. Concise, plain explainers on the day's politics, tech, and economics.
- 33.TLDR News Global · TLDR News. The global-affairs channel, for stories beyond the UK and US.
- 34.TLDR News EU · TLDR News. The European politics and policy channel from the same team.
Tools & Products
The AI products worth actually using, across code, writing, and media.
- 35.Cursor · Anysphere. The AI-native code editor most developers reach for first.
- 36.Claude Code · Anthropic. Anthropic's terminal coding agent, the pick for serious work.
- 37.Codex · OpenAI. OpenAI's coding agent, in the terminal, IDE, and the cloud.
- 38.Devin · Cognition. Autonomous software engineer that plans and ships code on its own.
- 39.v0 · Vercel. Generates working web apps and UI from a prompt.
- 40.Lovable · Lovable. Prompt-to-app builder for shipping full prototypes fast.
- 41.Base44 · Wix. No-code app builder that grew to a Wix acquisition in six months.
- 42.Perplexity · Perplexity AI. AI answer engine that cites its sources as it searches.
- 43.Exa · Exa AI. Search API that gives AI agents real-time, structured web data.
- 44.Granola · Granola. AI meeting notes that transcribe your calls without a bot joining.
- 45.NotebookLM · Google. Grounds answers in your own documents, and turns them into audio.
- 46.Whisper · OpenAI. The open speech-to-text model much of the ecosystem is built on.
- 47.ElevenLabs · ElevenLabs. The leading text-to-speech and voice-cloning platform.
- 48.Suno · Suno. Generates complete songs, vocals and all, from a text prompt.
- 49.Runway · Runway. AI video generation and editing tools for creative work.
- 50.Midjourney · Midjourney. Still the benchmark for aesthetic AI image generation.
Evals & Dev Tooling
Where to see who's actually ahead, and what to build and monitor with.
- 51.LMArena · LMArena (ex-LMSYS). Crowdsourced head-to-head voting that ranks models by human preference.
- 52.Artificial Analysis · Artificial Analysis. Independent leaderboards comparing models on intelligence, speed, and price.
- 53.SWE-bench · SWE-bench. The standard test of whether agents can fix real GitHub issues.
- 54.Epoch AI · Epoch AI. Research and data on AI trends, compute, and where the frontier is heading.
- 55.Hugging Face · Hugging Face. The hub for open models and datasets, and home to many leaderboards.
- 56.LangChain · LangChain. Agent framework (LangGraph) plus LangSmith for tracing and evals.
- 57.Langfuse · Langfuse. Open-source observability and evals for LLM apps. Now part of ClickHouse.
- 58.Weights & Biases · Weights & Biases. Experiment tracking and model tooling, widely used across ML teams.
People to Follow
Track these and you'll never fall too far behind. Follow a few on X and you'll feel the field move in real time.
- 59.Andrej Karpathy · @karpathy · Anthropic (ex-OpenAI, ex-Tesla). The field's best teacher. Read and watch everything he makes.
- 60.Sam Altman · @sama · OpenAI. OpenAI's CEO. Sets expectations for the field one post at a time.
- 61.Elon Musk · @elonmusk · xAI. Runs xAI, and X itself. Unavoidable, for better and worse.
- 62.Dario Amodei · @DarioAmodei · Anthropic. Anthropic's CEO. His long essays on scaling and safety set the terms of the debate.
- 63.Demis Hassabis · @demishassabis · Google DeepMind. Frontier research and the science applications, from AlphaFold to Gemini.
- 64.Greg Brockman · @gdb · OpenAI. OpenAI's president. Infrastructure scale, demos, and engineering culture.
- 65.Yann LeCun · @ylecun · Turing laureate. Turing laureate and the field's most reliable contrarian on LLMs.
- 66.Alexandr Wang · @alexandr_wang · Meta. Founded Scale AI, now Meta's Chief AI Officer leading its superintelligence push.
- 67.Jack Clark · Anthropic. His Import AI newsletter is the most trusted weekly read on where the field is going.
- 68.Lilian Weng · Ex-OpenAI. Deep, survey-grade posts that become the reference on their topic.
- 69.Simon Willison · Independent. The most reliable running commentary on what's actually new and useful.
- 70.Christopher Olah · Anthropic. Makes the inside of neural networks legible. Pioneer of interpretability.
Frontier AI Labs
Who's building the models, grouped by region and in no particular order. As of Aug 2026.
| Company | Product | Type | Remarks | Notable people | Weights | Founded | Region |
|---|---|---|---|---|---|---|---|
| OpenAI | GPT series · ChatGPT | Multimodal | Maker of ChatGPT and the GPT models. First to bring LLMs to a mass audience. | Sam Altman, Greg Brockman, Ilya Sutskever | Mostly closed | 2015 | United States |
| Anthropic | Claude series | Text + vision | Safety-focused research lab. Maker of the Claude models and the Claude Code agent. | Dario Amodei, Daniela Amodei, Jared Kaplan | Closed | 2021 | United States |
| Google DeepMind | Gemini series | Multimodal | Google's AI lab, maker of the Gemini models. Formed from DeepMind and Google Brain. | Demis Hassabis, Shane Legg, Koray Kavukcuoglu | Mostly closed | 2010 | United States |
| xAI | Grok series | Multimodal | Elon Musk's lab. Grok is integrated into X (formerly Twitter). | Elon Musk | Mixed | 2023 | United States |
| Meta AI | Llama series | Text + vision | Meta's AI lab. The Llama models are released as open weights. | Mark Zuckerberg, Alexandr Wang, Yann LeCun | Open weights | 2013 | United States |
| Microsoft AI | MAI series · Phi | Multimodal | Builds in-house MAI models and the open Phi family, alongside its OpenAI partnership. | Mustafa Suleyman, Karén Simonyan | Mixed | 2024 | United States |
| Safe Superintelligence | SSI (no product yet) | Undisclosed | Founded by Ilya Sutskever with a single stated goal. No product released yet. | Ilya Sutskever, Daniel Levy | Undisclosed | 2024 | United States |
| Thinking Machines Lab | Research + products | Undisclosed | Founded by ex-OpenAI CTO Mira Murati, with several ex-OpenAI researchers. | Mira Murati, John Schulman | Undisclosed | 2024 | United States |
| Reka AI | Reka series | Multimodal | Multimodal models from a team of ex-DeepMind and Meta researchers. | Yi Tay, Dani Yogatama | Mixed | 2022 | United States |
| Liquid AI | LFM series | Text + vision | MIT spin-out. Builds small, efficient models designed to run on-device. | Ramin Hasani, Mathias Lechner, Daniela Rus | Open weights | 2023 | United States |
| DeepSeek | DeepSeek series | Text + vision | Chinese lab known for open-weight models trained at relatively low cost. | Liang Wenfeng | Open weights | 2023 | China |
| Alibaba | Qwen series | Multimodal | Alibaba's model family. Wide range of open-weight sizes with broad language coverage. | Eddie Wu, Junyang Lin | Open weights | 2017 | China |
| Moonshot AI | Kimi series | Text + vision | Chinese lab behind Kimi, an early mover on long-context models. | Yang Zhilin | Open weights | 2023 | China |
| Zhipu AI | GLM series · Z.ai | Multimodal | Tsinghua University spin-out. Maker of the GLM open-weight models. | Tang Jie, Zhang Peng | Open weights | 2019 | China |
| MiniMax | MiniMax series | Multimodal (incl. video) | Chinese lab with multimodal models spanning text, audio, video, and music. | Yan Junjie | Open weights | 2021 | China |
| StepFun | Step series | Multimodal | Chinese lab focused on efficient, small active-parameter open-weight models. | Jiang Daxin, Zhang Xiangyu | Open weights | 2023 | China |
| ByteDance | Doubao / Seed series | Multimodal (incl. video) | TikTok's parent. Ships the Doubao consumer assistant and Seed research models. | Zhang Yiming | Mostly closed | 2023 | China |
| Baidu | ERNIE series | Multimodal | One of China's earliest movers on large models. Maker of the ERNIE / Wenxin family. | Robin Li | Mixed | 2013 | China |
| Tencent | Hunyuan series | Multimodal (incl. video) | Maker of the Hunyuan models, distributed across Tencent products including WeChat. | Pony Ma | Mixed | 2016 | China |
| Mistral AI | Mistral / Mixtral series | Text + vision | France-based lab. Maker of the Mistral and Mixtral open-weight models. | Arthur Mensch, Guillaume Lample, Timothée Lacroix | Mixed | 2023 | Europe |
Hyperscalers, Neoclouds & Compute
The clouds and chips the models actually run on.
| Company | Product | Category | Remarks | Founded | Region |
|---|---|---|---|---|---|
| Nvidia | GPUs · CUDA | Chipmaker | Designs the GPUs most models are trained on, and the CUDA software they run on. | 1993 | United States |
| Microsoft Azure | Azure AI · OpenAI partner | Hyperscaler | OpenAI's primary cloud partner. Serves models to enterprises through Azure AI. | 1975 | United States |
| Amazon Web Services | AWS · Bedrock | Hyperscaler | Largest cloud provider and an Anthropic backer. Bedrock hosts many models. | 1994 | United States |
| Google Cloud | Vertex AI · TPUs | Hyperscaler | Runs Gemini and rents its in-house TPUs, an alternative to training on Nvidia. | 1998 | United States |
| Oracle Cloud | OCI | Hyperscaler | Cloud provider with large GPU capacity deals supporting frontier training runs. | 1977 | United States |
| CoreWeave | GPU neocloud | Neocloud | Specialist GPU cloud that pivoted from crypto mining to renting AI compute. | 2017 | United States |
| Lambda | GPU cloud | Neocloud | GPU cloud for training and inference, one of the earliest AI-focused providers. | 2012 | United States |
| Together AI | Inference · fine-tuning | Neocloud | Runs and fine-tunes open models at scale, plus rents GPU clusters. | 2022 | United States |
| Fireworks AI | Inference platform | Neocloud | Fast, low-cost inference and training for open and custom models. | 2022 | United States |
| Crusoe | Energy-first GPU cloud | Neocloud | Builds and runs AI data centers powered by otherwise-wasted energy. | 2018 | United States |
| Nebius | AI cloud | Neocloud | Full-stack AI cloud for training and inference, spun out of Yandex. | 2024 | Europe |
| Cerebras | Wafer-scale chips · inference cloud | Chipmaker | Builds wafer-scale chips and runs them as a high-speed inference cloud. | 2015 | United States |
| AMD | Instinct GPUs · ROCm | Chipmaker | The main GPU alternative to Nvidia, with the Instinct line and ROCm software. | 1969 | United States |
| Broadcom | Custom AI chips · networking | Chipmaker | Co-designs custom accelerators for hyperscalers and supplies the networking between them. | 1991 | United States |
| TSMC | Chip foundry | Foundry | Fabricates nearly every leading AI chip. The largest contract chip manufacturer. | 1987 | Taiwan |
| ASML | EUV lithography | Equipment | Sole maker of the EUV lithography machines used to print the most advanced chips. | 1984 | Netherlands |
| SK Hynix | HBM memory | Memory | Leading supplier of the high-bandwidth memory (HBM) used in AI accelerators. | 1983 | South Korea |
| Samsung | HBM memory · foundry | Memory / Foundry | Makes HBM memory and runs a leading-edge foundry, competing with SK Hynix and TSMC. | 1969 | South Korea |
