Videos
| Thumb | Title | Channel | Status | Published |
|---|---|---|---|---|
|
|
DeepSeek V4 Flash Vision Is INSANE – Tested With DeepSeek Harness!
DeepSeek V4 Flash Vision (experimental) adds vision to a small model, and the results on coding tasks like browser... |
Bijan Bowen | summarized | 2026-08-24 12:36 |
|
|
Ox Alpha Is INSANE – Testing the Mysterious New Stealth Model!
Ox Alpha is a mysterious new reasoning model available on OpenRouter, and based on Bijan Bowen's tests, it produces... |
Bijan Bowen | summarized | 2026-08-21 12:42 |
|
|
Gemini 3.7 Flash Is HERE – Testing Google’s BEST Model Yet!
Google's Gemini 3.7 Flash is here, and it's faster, cheaper, and more capable than its predecessors. In tests, it... |
Bijan Bowen | summarized | 2026-08-17 13:54 |
|
|
GLM 5.3 Is HERE – Is THIS the BEST Open Model Yet?
GLM 5.3 is a 753B-parameter open-weight mixture-of-experts model that matches closed-source giants like Kimi K3 on... |
Bijan Bowen | summarized | 2026-08-16 09:01 |
|
|
AI News: ChatGPT Ultrafast, Grok 4.6, 3 New Open-Source Models, and more!
ChatGPT just got 14x faster thanks to Cerebras custom chips, making the model no longer the bottleneck—your computer... |
Matthew Berman | summarized | 2026-08-14 23:34 |
|
|
Qwen3.8 27B Is INSANE – This Is the BEST Local AI Model Yet!
Qwen 3.8 27B is a powerful local AI model that can run on a single GPU, showing impressive results in coding, game... |
Bijan Bowen | summarized | 2026-08-14 22:21 |
|
|
QWEN 3.8 27B Local AI Review
Qwen 3.8 27B is a huge leap over 3.6 for local AI agentic coding. In a zero-shot test, it built three playable retro... |
Digital Spaceport | summarized | 2026-08-14 21:10 |
|
|
Gemini 3.7 Flash Beats Sonnet 5 for $0.75 (But There’s a Catch)
Google's Gemini 3.7 Flash beats Claude Sonnet 5 on coding benchmarks at $0.75 per million tokens — half the price of... |
Cloud Codes | summarized | 2026-08-14 16:00 |
|
|
xAI's Real Plan to Win Isn't Grok 4.6
Grok 4.6 ties OpenAI's top model on a key benchmark at a fifth of the output price, but the model is the least... |
Cloud Codes | summarized | 2026-08-14 04:05 |
|
|
Grok 4.6 Is INSANE – Is THIS a Frontier Model?
Grok 4.6 is a serious competitor to models from OpenAI and Anthropic, especially in coding and front-end tasks. The... |
Bijan Bowen | summarized | 2026-08-13 11:50 |
|
|
DeepSeek V4 Pro Is HERE – Is THIS the BEST Open Model Yet?
DeepSeek V4 Pro (0813) is out of preview and it's a big leap over the undercooked preview version, but it's not the... |
Bijan Bowen | summarized | 2026-08-13 00:15 |
|
|
Best Local Coding Model Right Now? Meta Muse Glimmer Changes Everything
Meta released Muse Glimmer, a 30B-parameter open-source model (Apache 2.0) designed to fit on a 24GB GPU, built via... |
Cloud Codes | summarized | 2026-08-11 16:00 |
|
|
Ling 3.0 Tiny First Test – Can a Model THIS Small Really Code?
Ling 3.0 Tiny, a 7.9B parameter mixture-of-experts model with 1.3B active parameters, impresses in coding tests... |
Bijan Bowen | summarized | 2026-08-08 14:50 |
|
|
Qwen3.8 Max Is HERE – Is THIS the BEST Open Model Yet?
Alibaba released Qwen3.8 Max, a 2.4 trillion parameter mixture-of-experts model with 95 billion active parameters,... |
Bijan Bowen | summarized | 2026-08-03 13:25 |
|
|
I open-sourced my Agent Skills repo (it went viral)
David Ondrej open-sourced his personal repository of 42 agent skills for AI coding tools like Codex and Claude Code,... |
David Ondrej | summarized | 2026-08-02 18:39 |
|
|
GPT-5.6 Luna First Test – Hands-On With OpenAI’s CHEAPEST Model!
OpenAI's GPT-5.6 Luna is now its cheapest model after an 80% price cut, costing $0.20 per million input tokens and... |
Bijan Bowen | summarized | 2026-08-02 11:09 |
|
|
DeepSeek V4 Flash Is INSANE – The Best Small Model Yet!
DeepSeek V4 Flash is a newly released official version of a small but powerful AI model with 284B parameters (13B... |
Bijan Bowen | summarized | 2026-07-31 14:12 |
|
|
ThinkingCap - The Local Coding Model
Bottle Cap AI's ThinkingCap fine-tune of Qwen 3.6 27B reduces reasoning tokens by ~46% while preserving benchmark... |
Sam Witteveen | summarized | 2026-07-30 13:00 |
|
|
Ling 3.0 Flash First Test – A Surprisingly GOOD Coding Model!
Ling 3.0 Flash from Ant is a surprisingly competent coding model at 124B total/5.1B active parameters, outperforming... |
Bijan Bowen | summarized | 2026-07-28 13:18 |
|
|
DeepSWE: A Contamination-Resistant Coding Benchmark — James Shi, Datacurve
DeepSWE is a contamination-resistant coding benchmark from Datacurve, composed of 113 original software engineering... |
AI Engineer | summarized | 2026-07-26 18:10 |
|
|
I Tested Opus 5 vs. Fable 5. What You Need to Know.
Claude Opus 5 is often cheaper than Fable 5 and can outperform it on coding and verification tasks, but Fable 5... |
Nate Herk | summarized | 2026-07-24 23:38 |
|
|
Opus 5 is FINALLY here! (WOAH)
Anthropic's Claude Opus 5 was released, outperforming the larger Fable 5 on most benchmarks while costing about half... |
Matthew Berman | summarized | 2026-07-24 21:47 |
|
|
Claude Opus 5 is Going to Save You Money
Claude Opus 5 has been released with state-of-the-art performance on coding and knowledge work benchmarks, often... |
Nate Herk | summarized | 2026-07-24 17:45 |
|
|
Is Kimi K3 Really That Good?! (Don't Just Believe The Hype)
Kimi K3 is the most powerful open-weight model released, but it suffers from reliability issues that public... |
Cole Medin | summarized | 2026-07-24 14:00 |
|
|
Poolside Laguna S2.1 First Test – A VERY Creative Local Model!
Poolside Laguna S2.1 is a 118B parameter Mixture of Experts model (8B active) that excels at creative writing and... |
Bijan Bowen | summarized | 2026-07-24 12:15 |
|
|
Build Anything with Kimi K3, Here’s How
Kimi K3 is an open-source AI model from Moonshot AI that matches or beats closed-source models like Fable 5 and... |
David Ondrej | summarized | 2026-07-24 09:12 |
|
|
Poolside’s Model Factory, Laguna S, Open Models, and the Race to AGI — Eiso Kant, Poolside AI
Poolside's co-founder Eiso Kant argues that model building is primarily an engineering discipline, and that their... |
Latent Space | summarized | 2026-07-22 20:46 |
|
|
Gemini 3.6 Flash Is HERE – Testing Google’s BEST Model Yet!
Google's Gemini 3.6 Flash is a new model that offers significant speed improvements, a 17% reduction in token usage,... |
Bijan Bowen | summarized | 2026-07-21 21:06 |
|
|
Kimi K3 Is INSANE – Is THIS a Sol & Fable Competitor?
Kimi K3 is a massive 2.8 trillion parameter open-weight model that claims to rival proprietary models like Claude... |
Bijan Bowen | summarized | 2026-07-16 20:52 |
|
|
Meta Muse Spark 1.1 First Test – Is THIS a Frontier Model?
Meta's Muse Spark 1.1 shows benchmark improvements over its predecessor but fails in practical coding and game... |
Bijan Bowen | summarized | 2026-07-11 16:41 |
|
|
I Tested GPT 5.6 Sol vs Fable (4 Real Uses Cases)
GPT 5.6 Soul outperforms Fable 5 in three out of four real-world coding and design tasks at a fraction of the cost,... |
Jason Lee | summarized | 2026-07-11 10:32 |
|
|
GPT-5.6 Is HERE – Is THIS the Most INFURIATING Model Yet?
OpenAI released GPT-5.6 as a trio of models—Soul, Terra, and Luna—each with different cost and capability tiers,... |
Bijan Bowen | summarized | 2026-07-10 01:43 |
|
|
Grok 4.5 explained in 8min..
Grok 4.5 is a 1.5-trillion-parameter frontier model from xAI, built on a new V9 foundation that incorporates... |
Caleb Writes Code | summarized | 2026-07-10 01:04 |
|
|
Grok 4.5 Is INSANE – Is THIS a GPT & Opus Competitor?
Grok 4.5 is a highly competitive AI model that rivals state-of-the-art models like GPT-5 and Opus 4, offering strong... |
Bijan Bowen | summarized | 2026-07-09 04:03 |
|
|
Field Guide to Fable — Thariq Shihipar, Anthropic
Anthropic's Thariq Shihipar introduces Fable, a new Claude model that represents a major leap in capability, akin to... |
AI Engineer | summarized | 2026-07-06 16:00 |
|
|
Fable 5 is back..
Fable 5 is Anthropic's latest high-intelligence model, released after a 17-day suspension for safety overhauls, and... |
Caleb Writes Code | summarized | 2026-07-03 00:52 |
|
|
Fable 5 is back… here is my plan
Fable 5 is back after a brief ban, and the speaker considers it a step-change in model capability, especially for... |
David Ondrej | summarized | 2026-07-02 19:50 |
|
|
Why the Next 12 Months Will Create More AI Wealth Than the Last 100 Years
Three major AI companies (SpaceX, Anthropic, OpenAI) are filing for IPOs at a combined ~$4 trillion valuation,... |
AI Founders | summarized | 2026-07-02 16:00 |
|
|
Claude Sonnet 5 Is HERE – Hands-On With Anthropic’s NEW Model!
Anthropic released Claude Sonnet 5 alongside a Linux desktop app beta. The model offers improved agentic coding and... |
Bijan Bowen | summarized | 2026-06-30 23:29 |
|
|
GLM-5.2 vs Claude Opus 4.8 – Does GLM REALLY Beat Claude?
GLM 5.2, an open-weights model, is compared head-to-head against Claude Opus 4.8 across several challenging tasks: a... |
Bijan Bowen | summarized | 2026-06-29 15:57 |
|
|
Ornith 1.0 First Look & Test – The BEST New Local Coding Models?
Ornith 1.0, a family of fine-tuned local coding models (9B dense and 35B MoE variants tested), shows promising... |
Bijan Bowen | summarized | 2026-06-28 11:11 |
|
|
Introducing Ornith 1.0 - Agentic Coding LLMs
Ornith 1.0 from Deep Reinforce introduces self-scaffolding LLMs for agentic coding, where the model learns to... |
Sam Witteveen | summarized | 2026-06-26 14:00 |
|
|
I Tested the Fable 5 Killer (Hermes Agent)
Fable 5 is not actually killed by Sakana Fugu or GLM 5.2; the tests show that Fugu is an intelligent router that... |
Jack Roberts | summarized | 2026-06-24 16:59 |
|
|
North Mini Code First Test – Cohere’s LOCAL Agentic Coding Model!
Cohere's North Mini Code is a 30-billion parameter mixture-of-experts Apache 2.0 licensed coding model optimized for... |
Bijan Bowen | summarized | 2026-06-20 13:20 |
|
|
OpenRouter Fusion First Test – Does THIS Beat Claude Fable?
OpenRouter Fusion, which synthesizes outputs from multiple models via a judge model, shows strong performance on... |
Bijan Bowen | summarized | 2026-06-16 13:20 |
|
|
Omnigent: The New Meta-Harness for EVERY Coding Agent - Claude Code, Codex, Pi, More
OmniAgent is a new open-source meta harness from Databricks that orchestrates multiple AI coding assistants like... |
Cole Medin | summarized | 2026-06-15 14:42 |
|
|
MYTHOS MYTHOS MYTHOS
Anthropic released Mythos 5 and Fable 5, a new class of 10-trillion-parameter models that exceed all previous models... |
Matthew Berman | summarized | 2026-06-09 23:02 |
|
|
Anthropic Just Dropped Claude Mythos and Fable 5 (Full Breakdown)
Anthropic released Claude Fable 5 and Claude Mythos 5, with Fable 5 being a safer, public version of the previously... |
Brock Mesarich | AI for Non Techies | summarized | 2026-06-09 19:35 |
|
|
Opus 4.8 Just Dropped. Here's How To Actually Use It.
Opus 4.8 is released with improved benchmarks and features, including sharper judgment, more honesty, and the... |
Nate Herk | summarized | 2026-05-28 18:52 |
|
|
Finally a good benchmark (DeepSWE)
Deep Suite is a new benchmark for coding models that delivers four major advances over today's public benchmarks,... |
Matthew Berman | summarized | 2026-05-27 16:03 |
Frontier News · by Hyperjump Technology