Resources
Sadly, I am only human, and my context window is embarrassingly small. There is too much happening in AI to remember everything I read and learn. So this is my external memory, an AI info-dump for me and anyone else trying to keep up.
Foundations
The explainers that make everything after them click.
- 1.Neural Networks: Zero to Hero (opens in new tab) · Andrej Karpathy. Builds a neural net, then a GPT, from scratch in code. The single best on-ramp there is.
- 2.The Illustrated Transformer (opens in new tab) · Jay Alammar. The picture-first walkthrough of attention that most people first understood transformers from.
- 3.The Unreasonable Effectiveness of Recurrent Neural Networks (opens in new tab) · Andrej Karpathy, 2015. The post that made a generation of engineers fall for sequence models.
- 4.Understanding LSTM Networks (opens in new tab) · Christopher Olah, 2015. The canonical, diagram-driven explanation of how gated recurrent memory works.
Landmark Papers
Must reads.
- 5.Attention Is All You Need (opens in new tab) · Ashish Vaswani et al., 2017. Introduced the transformer. The paper the modern era is built on.
- 6.Efficient Estimation of Word Representations in Vector Space (opens in new tab) · Tomas Mikolov et al. (Google), 2013. Word2Vec. Showed words could be turned into meaningful vectors, the root of embeddings.
- 7.Improving Language Understanding by Generative Pre-Training (opens in new tab) · Alec Radford et al. (OpenAI), 2018. GPT-1. First showed pre-train then fine-tune on a transformer.
- 8.BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (opens in new tab) · Jacob Devlin et al. (Google), 2018. The bidirectional encoder that dominated NLP benchmarks before the GPT line pulled ahead.
- 9.Language Models are Unsupervised Multitask Learners (opens in new tab) · Alec Radford et al. (OpenAI), 2019. GPT-2. Scaled the idea up and hinted at zero-shot ability.
- 10.Language Models are Few-Shot Learners (opens in new tab) · Tom Brown et al. (OpenAI), 2020. GPT-3. The one that showed scale alone unlocks new behaviour.
- 11.ImageNet Classification with Deep Convolutional Neural Networks (opens in new tab) · Alex Krizhevsky, Ilya Sutskever, Geoffrey Hinton, 2012. AlexNet. The result that set off the modern deep-learning era.
- 12.Sequence to Sequence Learning with Neural Networks (opens in new tab) · Ilya Sutskever, Oriol Vinyals, Quoc Le, 2014. Seq2Seq. The encoder-decoder idea that set the path toward modern language models.
- 13.Deep Residual Learning for Image Recognition (opens in new tab) · Kaiming He et al., 2015. ResNet. Made truly deep networks trainable and reset the field.
- 14.Training Language Models to Follow Instructions with Human Feedback (opens in new tab) · Long Ouyang et al. (OpenAI), 2022. InstructGPT. The RLHF recipe that turned raw models into assistants.
Courses
If you want to go from reading to doing.
- 15.Practical Deep Learning for Coders (opens in new tab) · fast.ai. Top-down and hands-on. Ship a working model in the first lesson.
- 16.CS231n: Deep Learning for Computer Vision (opens in new tab) · Stanford. The course Karpathy designed. Still the gold standard for the fundamentals.
- 17.CS224n: NLP with Deep Learning (opens in new tab) · Stanford. The rigorous path through language models, from embeddings up.
- 18.Machine Learning Specialization (opens in new tab) · Andrew Ng · DeepLearning.AI. The gentlest on-ramp to the maths and intuitions underneath it all.
Watch & Listen
Talks, channels, and podcasts. Sit back and let someone explain it.
- 19.Neural Networks (series) (opens in new tab) · 3Blue1Brown. The most beautiful visual intuition for what a network actually does.
- 20.[1hr Talk] Intro to Large Language Models (opens in new tab) · Andrej Karpathy, 2023. The clearest one-hour mental model of how LLMs work and where they go.
- 21.Let's build GPT: from scratch, in code (opens in new tab) · Andrej Karpathy. Watch a GPT get built line by line. Everything demystifies at once.
- 22.AI Explained (opens in new tab) · YouTube channel. Calm, careful breakdowns of every major model release and benchmark.
- 23.Theo, t3.gg (opens in new tab) · YouTube channel. Opinionated, fast-moving takes on AI coding tools and how they change the work.
- 24.Prof G Markets (opens in new tab) · Scott Galloway · YouTube channel. The business and money side of the AI boom, markets, valuations, and who's winning.
- 25.YC Paper Club (opens in new tab) · Y Combinator. Researchers and founders break down the latest AI papers. A great way to keep up.
- 26.AI Engineer (opens in new tab) · swyx & the AI Engineer team. Talks from the AI Engineer conferences. Practitioners showing what they actually shipped.
- 27.Dwarkesh Podcast (opens in new tab) · Dwarkesh Patel. Deeply researched, unusually hard questions for the people building frontier AI.
- 28.Latent Space (opens in new tab) · Shawn Wang (swyx) & Alessio Fanelli. The AI engineer's trade publication in audio form. Real implementation, not hype.
- 29.No Priors (opens in new tab) · Sarah Guo & Elad Gil. Frontier AI from the investor and product-builder angle. Where the field is heading.
- 30.Lex Fridman Podcast (opens in new tab) · Lex Fridman. Long, unhurried conversations with many of the field's defining figures.
- 31.The 80,000 Hours Podcast (opens in new tab) · Rob Wiblin & team. In-depth interviews on AI safety and the field's biggest risks, now on video.
- 32.TLDR News (opens in new tab) · TLDR News. Concise, plain explainers on the day's politics, tech, and economics.
- 33.TLDR News Global (opens in new tab) · TLDR News. The global-affairs channel, for stories beyond the UK and US.
- 34.TLDR News EU (opens in new tab) · TLDR News. The European politics and policy channel from the same team.
Tools & Products
The AI products worth actually using, across code, writing, and media.
- 35.Cursor (opens in new tab) · Anysphere. The AI-native code editor most developers reach for first.
- 36.Claude Code (opens in new tab) · Anthropic. Anthropic's terminal coding agent, the pick for serious work.
- 37.Codex (opens in new tab) · OpenAI. OpenAI's coding agent, in the terminal, IDE, and the cloud.
- 38.Devin (opens in new tab) · Cognition. Autonomous software engineer that plans and ships code on its own.
- 39.v0 (opens in new tab) · Vercel. Generates working web apps and UI from a prompt.
- 40.Lovable (opens in new tab) · Lovable. Prompt-to-app builder for shipping full prototypes fast.
- 41.Base44 (opens in new tab) · Wix. No-code app builder that grew to a Wix acquisition in six months.
- 42.Perplexity (opens in new tab) · Perplexity AI. AI answer engine that cites its sources as it searches.
- 43.Exa (opens in new tab) · Exa AI. Search API that gives AI agents real-time, structured web data.
- 44.Granola (opens in new tab) · Granola. AI meeting notes that transcribe your calls without a bot joining.
- 45.NotebookLM (opens in new tab) · Google. Grounds answers in your own documents, and turns them into audio.
- 46.Whisper (opens in new tab) · OpenAI. The open speech-to-text model much of the ecosystem is built on.
- 47.ElevenLabs (opens in new tab) · ElevenLabs. The leading text-to-speech and voice-cloning platform.
- 48.Suno (opens in new tab) · Suno. Generates complete songs, vocals and all, from a text prompt.
- 49.Runway (opens in new tab) · Runway. AI video generation and editing tools for creative work.
- 50.Midjourney (opens in new tab) · Midjourney. Still the benchmark for aesthetic AI image generation.
Evals & Dev Tooling
Where to see who's actually ahead, and what to build and monitor with.
- 51.LMArena (opens in new tab) · LMArena (ex-LMSYS). Crowdsourced head-to-head voting that ranks models by human preference.
- 52.Artificial Analysis (opens in new tab) · Artificial Analysis. Independent leaderboards comparing models on intelligence, speed, and price.
- 53.SWE-bench (opens in new tab) · SWE-bench. The standard test of whether agents can fix real GitHub issues.
- 54.Epoch AI (opens in new tab) · Epoch AI. Research and data on AI trends, compute, and where the frontier is heading.
- 55.Hugging Face (opens in new tab) · Hugging Face. The hub for open models and datasets, and home to many leaderboards.
- 56.LangChain (opens in new tab) · LangChain. Agent framework (LangGraph) plus LangSmith for tracing and evals.
- 57.Langfuse (opens in new tab) · Langfuse. Open-source observability and evals for LLM apps. Now part of ClickHouse.
- 58.Weights & Biases (opens in new tab) · Weights & Biases. Experiment tracking and model tooling, widely used across ML teams.
People to Follow
Track these and you'll never fall too far behind. Follow a few on X and you'll feel the field move in real time.
- 59.Andrej Karpathy (opens in new tab) · @karpathy · Anthropic (ex-OpenAI, ex-Tesla). The field's best teacher. Read and watch everything he makes.
- 60.Sam Altman (opens in new tab) · @sama · OpenAI. OpenAI's CEO. Sets expectations for the field one post at a time.
- 61.Elon Musk (opens in new tab) · @elonmusk · xAI. Runs xAI, and X itself. Unavoidable, for better and worse.
- 62.Dario Amodei (opens in new tab) · @DarioAmodei · Anthropic. Anthropic's CEO. His long essays on scaling and safety set the terms of the debate.
- 63.Demis Hassabis (opens in new tab) · @demishassabis · Google DeepMind. Frontier research and the science applications, from AlphaFold to Gemini.
- 64.Greg Brockman (opens in new tab) · @gdb · OpenAI. OpenAI's president. Infrastructure scale, demos, and engineering culture.
- 65.Yann LeCun (opens in new tab) · @ylecun · Turing laureate. Turing laureate and the field's most reliable contrarian on LLMs.
- 66.Alexandr Wang (opens in new tab) · @alexandr_wang · Meta. Founded Scale AI, now Meta's Chief AI Officer leading its superintelligence push.
- 67.Jack Clark (opens in new tab) · Anthropic. His Import AI newsletter is the most trusted weekly read on where the field is going.
- 68.Lilian Weng (opens in new tab) · Ex-OpenAI. Deep, survey-grade posts that become the reference on their topic.
- 69.Simon Willison (opens in new tab) · Independent. The most reliable running commentary on what's actually new and useful.
- 70.Christopher Olah (opens in new tab) · Anthropic. Makes the inside of neural networks legible. Pioneer of interpretability.
Frontier AI Labs
Who's building the models, grouped by region and in no particular order. As of Aug 2026.
| Company | Product | Type | Remarks | Notable people | Weights | Founded | Region |
|---|---|---|---|---|---|---|---|
| OpenAI (opens in new tab) | GPT series · ChatGPT | Multimodal | Maker of ChatGPT and the GPT models. First to bring LLMs to a mass audience. | Sam Altman, Greg Brockman, Ilya Sutskever | Mostly closed | 2015 | United States |
| Anthropic (opens in new tab) | Claude series | Text + vision | Safety-focused research lab. Maker of the Claude models and the Claude Code agent. | Dario Amodei, Daniela Amodei, Jared Kaplan | Closed | 2021 | United States |
| Google DeepMind (opens in new tab) | Gemini series | Multimodal | Google's AI lab, maker of the Gemini models. Formed from DeepMind and Google Brain. | Demis Hassabis, Shane Legg, Koray Kavukcuoglu | Mostly closed | 2010 | United States |
| xAI (opens in new tab) | Grok series | Multimodal | Elon Musk's lab. Grok is integrated into X (formerly Twitter). | Elon Musk | Mixed | 2023 | United States |
| Meta AI (opens in new tab) | Llama series | Text + vision | Meta's AI lab. The Llama models are released as open weights. | Mark Zuckerberg, Alexandr Wang, Yann LeCun | Open weights | 2013 | United States |
| Microsoft AI (opens in new tab) | MAI series · Phi | Multimodal | Builds in-house MAI models and the open Phi family, alongside its OpenAI partnership. | Mustafa Suleyman, Karén Simonyan | Mixed | 2024 | United States |
| Safe Superintelligence (opens in new tab) | SSI (no product yet) | Undisclosed | Founded by Ilya Sutskever with a single stated goal. No product released yet. | Ilya Sutskever, Daniel Levy | Undisclosed | 2024 | United States |
| Thinking Machines Lab (opens in new tab) | Research + products | Undisclosed | Founded by ex-OpenAI CTO Mira Murati, with several ex-OpenAI researchers. | Mira Murati, John Schulman | Undisclosed | 2024 | United States |
| Reka AI (opens in new tab) | Reka series | Multimodal | Multimodal models from a team of ex-DeepMind and Meta researchers. | Yi Tay, Dani Yogatama | Mixed | 2022 | United States |
| Liquid AI (opens in new tab) | LFM series | Text + vision | MIT spin-out. Builds small, efficient models designed to run on-device. | Ramin Hasani, Mathias Lechner, Daniela Rus | Open weights | 2023 | United States |
| DeepSeek (opens in new tab) | DeepSeek series | Text + vision | Chinese lab known for open-weight models trained at relatively low cost. | Liang Wenfeng | Open weights | 2023 | China |
| Alibaba (opens in new tab) | Qwen series | Multimodal | Alibaba's model family. Wide range of open-weight sizes with broad language coverage. | Eddie Wu, Junyang Lin | Open weights | 2017 | China |
| Moonshot AI (opens in new tab) | Kimi series | Text + vision | Chinese lab behind Kimi, an early mover on long-context models. | Yang Zhilin | Open weights | 2023 | China |
| Zhipu AI (opens in new tab) | GLM series · Z.ai | Multimodal | Tsinghua University spin-out. Maker of the GLM open-weight models. | Tang Jie, Zhang Peng | Open weights | 2019 | China |
| MiniMax (opens in new tab) | MiniMax series | Multimodal (incl. video) | Chinese lab with multimodal models spanning text, audio, video, and music. | Yan Junjie | Open weights | 2021 | China |
| StepFun (opens in new tab) | Step series | Multimodal | Chinese lab focused on efficient, small active-parameter open-weight models. | Jiang Daxin, Zhang Xiangyu | Open weights | 2023 | China |
| ByteDance (opens in new tab) | Doubao / Seed series | Multimodal (incl. video) | TikTok's parent. Ships the Doubao consumer assistant and Seed research models. | Zhang Yiming | Mostly closed | 2023 | China |
| Baidu (opens in new tab) | ERNIE series | Multimodal | One of China's earliest movers on large models. Maker of the ERNIE / Wenxin family. | Robin Li | Mixed | 2013 | China |
| Tencent (opens in new tab) | Hunyuan series | Multimodal (incl. video) | Maker of the Hunyuan models, distributed across Tencent products including WeChat. | Pony Ma | Mixed | 2016 | China |
| Mistral AI (opens in new tab) | Mistral / Mixtral series | Text + vision | France-based lab. Maker of the Mistral and Mixtral open-weight models. | Arthur Mensch, Guillaume Lample, Timothée Lacroix | Mixed | 2023 | Europe |
Hyperscalers, Neoclouds & Compute
The clouds and chips the models actually run on.
| Company | Product | Category | Remarks | Founded | Region |
|---|---|---|---|---|---|
| Nvidia (opens in new tab) | GPUs · CUDA | Chipmaker | Designs the GPUs most models are trained on, and the CUDA software they run on. | 1993 | United States |
| Microsoft Azure (opens in new tab) | Azure AI · OpenAI partner | Hyperscaler | OpenAI's primary cloud partner. Serves models to enterprises through Azure AI. | 1975 | United States |
| Amazon Web Services (opens in new tab) | AWS · Bedrock | Hyperscaler | Largest cloud provider and an Anthropic backer. Bedrock hosts many models. | 1994 | United States |
| Google Cloud (opens in new tab) | Vertex AI · TPUs | Hyperscaler | Runs Gemini and rents its in-house TPUs, an alternative to training on Nvidia. | 1998 | United States |
| Oracle Cloud (opens in new tab) | OCI | Hyperscaler | Cloud provider with large GPU capacity deals supporting frontier training runs. | 1977 | United States |
| CoreWeave (opens in new tab) | GPU neocloud | Neocloud | Specialist GPU cloud that pivoted from crypto mining to renting AI compute. | 2017 | United States |
| Lambda (opens in new tab) | GPU cloud | Neocloud | GPU cloud for training and inference, one of the earliest AI-focused providers. | 2012 | United States |
| Together AI (opens in new tab) | Inference · fine-tuning | Neocloud | Runs and fine-tunes open models at scale, plus rents GPU clusters. | 2022 | United States |
| Fireworks AI (opens in new tab) | Inference platform | Neocloud | Fast, low-cost inference and training for open and custom models. | 2022 | United States |
| Crusoe (opens in new tab) | Energy-first GPU cloud | Neocloud | Builds and runs AI data centers powered by otherwise-wasted energy. | 2018 | United States |
| Nebius (opens in new tab) | AI cloud | Neocloud | Full-stack AI cloud for training and inference, spun out of Yandex. | 2024 | Europe |
| Cerebras (opens in new tab) | Wafer-scale chips · inference cloud | Chipmaker | Builds wafer-scale chips and runs them as a high-speed inference cloud. | 2015 | United States |
| AMD (opens in new tab) | Instinct GPUs · ROCm | Chipmaker | The main GPU alternative to Nvidia, with the Instinct line and ROCm software. | 1969 | United States |
| Broadcom (opens in new tab) | Custom AI chips · networking | Chipmaker | Co-designs custom accelerators for hyperscalers and supplies the networking between them. | 1991 | United States |
| TSMC (opens in new tab) | Chip foundry | Foundry | Fabricates nearly every leading AI chip. The largest contract chip manufacturer. | 1987 | Taiwan |
| ASML (opens in new tab) | EUV lithography | Equipment | Sole maker of the EUV lithography machines used to print the most advanced chips. | 1984 | Netherlands |
| SK Hynix (opens in new tab) | HBM memory | Memory | Leading supplier of the high-bandwidth memory (HBM) used in AI accelerators. | 1983 | South Korea |
| Samsung (opens in new tab) | HBM memory · foundry | Memory / Foundry | Makes HBM memory and runs a leading-edge foundry, competing with SK Hynix and TSMC. | 1969 | South Korea |
