Episodes (Page 8)
✨
Drew Houston has spent 400+ hours coding with LLMs and is refocusing Dropbox's 2,500+ employees around AI-native development 17 years after founding the company
✨
Ankur Goyal of Braintrust argues that production AI engineering should start with evals, not operational tooling, following the pattern of successful LLMOps founders with AI/research backgrounds
✨
OpenAI DevDay 2024 focused on developer-facing API announcements including Realtime API, Vision Finetuning, Prompt Caching, and Model Distillation rather than ChatGPT product announcements
✨
OpenAI's o1 release and recent hiring of Noam Brown and Shunyu Yao signals focus on tool-using chain-of-thought and tree-of-thought architectures for Level 3 Agents
✨
Sander Schulhoff's 'The Prompt Report' synthesizes 1,600+ arXiv papers on prompting techniques including few-shot learning, chain-of-thought, tree search, and self-criticism strategies
✨
Michelle Pokrass and OpenAI's DevRel team cover the entire OpenAI product suite including ChatGPT-latest, GPT-4o, o1 models, and how they're delivered via API with Structured Outputs
Michelle Pokrass
✨
AI inference costs decreased 10-100x in 2024, with open models like Llama 3.1 405B costing $3/mtok versus $30/mtok for Claude 3 Opus, and frontier models dropped 400x from 2022-2024
✨
Nicholas Carlini's 'How I Use AI' blog post demonstrates a practical approach focused on individual AI applications rather than broad AGI potential, covering 12 use cases with specific prompts
Nicholas Carlini
✨
Cosine Genie achieved #1 ranking on SWE-Bench Full, Lite, and Verified using GPT-4o fine-tuning at scale on billions of tokens of synthetic data, beating all other agents including Cognition's Devin
Alistair Pullen
✨
Jeremy Howard's Answer.AI ships 1000s of successful AI products with no managers and a team of 12, focusing on practical AI R&D aligned with GPU-poor needs
✨
Meta's Segment Anything 2 (SAM 2) improves image segmentation accuracy while being 6x faster than SAM 1, and elegantly solved video segmentation with 3x fewer interactions than prior approaches
✨
Q2 2024 AI progress analyzed through Four Wars framework: GPU-rich frontier labs (Claude 3.5, Mistral Large), GPU-rich helping GPU-poors (Llama 3.1 synthetic data, Phi 3, Gemma 2), and on-device LL...
✨
Meta released Llama 3.1-405B, the largest open source model trained on 15T tokens beating GPT-4 on benchmarks, with 8B and 70B models also receiving significant spec bumps
✨
Clémentine Fourrier leads HuggingFace's OpenLLM Leaderboard, which standardizes model evaluation using high-quality benchmarks with reproducible, centralized scoring to replace lab-specific reports
✨
Reka AI achieved #7 on LMsys leaderboard with only 20 employees and $60M funding, demonstrating that top-tier model performance no longer requires massive teams like OpenAI (600) or Google (950+ co...
✨
Databricks' DBRX and Imbue's 70B model outperform GPT-4o zero-shot on reasoning/coding benchmarks while using 7x less data than Llama 3 70B
✨
Raza Habib of HumanLoop hosts High Agency podcast, flipping the interview dynamic with Shawn Wang to discuss the AI Engineer World's Fair and the relevance of the 'Rise of the AI Engineer' essay on...
✨
James Brady (Head of Engineering) and Adam Wiggins (Cofounder Ink & Switch, Heroku) from Elicit share hiring strategies for AI engineers, defining the role as conventional engineers with LLM and pr...
James Brady
✨
Mike Conover, who led OSS models at Databricks and created Dolly, founded Brightwave as an AI research assistant for investment professionals and announced $6M seed round led by Alessio and Decibel
✨
Discusses code editing benchmarks (WebArena, Sotopia), OpenDevin agent framework, and tensions between academic research and industry implementation of AI systems
Aman Sanger
Graham Neubig
Moritz Hardt