Videos list
Thumb Title Channel Status Published
DeepSeek V4 Flash Vision Is INSANE – Tested With DeepSeek Harness!

DeepSeek V4 Flash Vision (experimental) adds vision to a small model, and the results on coding tasks like browser...

Bijan Bowen summarized 2026-08-24 12:36
Ox Alpha Is INSANE – Testing the Mysterious New Stealth Model!

Ox Alpha is a mysterious new reasoning model available on OpenRouter, and based on Bijan Bowen's tests, it produces...

Bijan Bowen summarized 2026-08-21 12:42
Gemini 3.7 Flash Is HERE – Testing Google’s BEST Model Yet!

Google's Gemini 3.7 Flash is here, and it's faster, cheaper, and more capable than its predecessors. In tests, it...

Bijan Bowen summarized 2026-08-17 13:54
GLM 5.3 Is HERE – Is THIS the BEST Open Model Yet?

GLM 5.3 is a 753B-parameter open-weight mixture-of-experts model that matches closed-source giants like Kimi K3 on...

Bijan Bowen summarized 2026-08-16 09:01
AI News: ChatGPT Ultrafast, Grok 4.6, 3 New Open-Source Models, and more!

ChatGPT just got 14x faster thanks to Cerebras custom chips, making the model no longer the bottleneck—your computer...

Matthew Berman summarized 2026-08-14 23:34
Qwen3.8 27B Is INSANE – This Is the BEST Local AI Model Yet!

Qwen 3.8 27B is a powerful local AI model that can run on a single GPU, showing impressive results in coding, game...

Bijan Bowen summarized 2026-08-14 22:21
QWEN 3.8 27B Local AI Review

Qwen 3.8 27B is a huge leap over 3.6 for local AI agentic coding. In a zero-shot test, it built three playable retro...

Digital Spaceport summarized 2026-08-14 21:10
Gemini 3.7 Flash Beats Sonnet 5 for $0.75 (But There’s a Catch)

Google's Gemini 3.7 Flash beats Claude Sonnet 5 on coding benchmarks at $0.75 per million tokens — half the price of...

Cloud Codes summarized 2026-08-14 16:00
xAI's Real Plan to Win Isn't Grok 4.6

Grok 4.6 ties OpenAI's top model on a key benchmark at a fifth of the output price, but the model is the least...

Cloud Codes summarized 2026-08-14 04:05
Grok 4.6 Is INSANE – Is THIS a Frontier Model?

Grok 4.6 is a serious competitor to models from OpenAI and Anthropic, especially in coding and front-end tasks. The...

Bijan Bowen summarized 2026-08-13 11:50
DeepSeek V4 Pro Is HERE – Is THIS the BEST Open Model Yet?

DeepSeek V4 Pro (0813) is out of preview and it's a big leap over the undercooked preview version, but it's not the...

Bijan Bowen summarized 2026-08-13 00:15
Best Local Coding Model Right Now? Meta Muse Glimmer Changes Everything

Meta released Muse Glimmer, a 30B-parameter open-source model (Apache 2.0) designed to fit on a 24GB GPU, built via...

Cloud Codes summarized 2026-08-11 16:00
Ling 3.0 Tiny First Test – Can a Model THIS Small Really Code?

Ling 3.0 Tiny, a 7.9B parameter mixture-of-experts model with 1.3B active parameters, impresses in coding tests...

Bijan Bowen summarized 2026-08-08 14:50
Qwen3.8 Max Is HERE – Is THIS the BEST Open Model Yet?

Alibaba released Qwen3.8 Max, a 2.4 trillion parameter mixture-of-experts model with 95 billion active parameters,...

Bijan Bowen summarized 2026-08-03 13:25
I open-sourced my Agent Skills repo (it went viral)

David Ondrej open-sourced his personal repository of 42 agent skills for AI coding tools like Codex and Claude Code,...

David Ondrej summarized 2026-08-02 18:39
GPT-5.6 Luna First Test – Hands-On With OpenAI’s CHEAPEST Model!

OpenAI's GPT-5.6 Luna is now its cheapest model after an 80% price cut, costing $0.20 per million input tokens and...

Bijan Bowen summarized 2026-08-02 11:09
DeepSeek V4 Flash Is INSANE – The Best Small Model Yet!

DeepSeek V4 Flash is a newly released official version of a small but powerful AI model with 284B parameters (13B...

Bijan Bowen summarized 2026-07-31 14:12
ThinkingCap - The Local Coding Model

Bottle Cap AI's ThinkingCap fine-tune of Qwen 3.6 27B reduces reasoning tokens by ~46% while preserving benchmark...

Sam Witteveen summarized 2026-07-30 13:00
Ling 3.0 Flash First Test – A Surprisingly GOOD Coding Model!

Ling 3.0 Flash from Ant is a surprisingly competent coding model at 124B total/5.1B active parameters, outperforming...

Bijan Bowen summarized 2026-07-28 13:18
DeepSWE: A Contamination-Resistant Coding Benchmark — James Shi, Datacurve

DeepSWE is a contamination-resistant coding benchmark from Datacurve, composed of 113 original software engineering...

AI Engineer summarized 2026-07-26 18:10
I Tested Opus 5 vs. Fable 5. What You Need to Know.

Claude Opus 5 is often cheaper than Fable 5 and can outperform it on coding and verification tasks, but Fable 5...

Nate Herk summarized 2026-07-24 23:38
Opus 5 is FINALLY here! (WOAH)

Anthropic's Claude Opus 5 was released, outperforming the larger Fable 5 on most benchmarks while costing about half...

Matthew Berman summarized 2026-07-24 21:47
Claude Opus 5 is Going to Save You Money

Claude Opus 5 has been released with state-of-the-art performance on coding and knowledge work benchmarks, often...

Nate Herk summarized 2026-07-24 17:45
Is Kimi K3 Really That Good?! (Don't Just Believe The Hype)

Kimi K3 is the most powerful open-weight model released, but it suffers from reliability issues that public...

Cole Medin summarized 2026-07-24 14:00
Poolside Laguna S2.1 First Test – A VERY Creative Local Model!

Poolside Laguna S2.1 is a 118B parameter Mixture of Experts model (8B active) that excels at creative writing and...

Bijan Bowen summarized 2026-07-24 12:15
Build Anything with Kimi K3, Here’s How

Kimi K3 is an open-source AI model from Moonshot AI that matches or beats closed-source models like Fable 5 and...

David Ondrej summarized 2026-07-24 09:12
Poolside’s Model Factory, Laguna S, Open Models, and the Race to AGI — Eiso Kant, Poolside AI

Poolside's co-founder Eiso Kant argues that model building is primarily an engineering discipline, and that their...

Latent Space summarized 2026-07-22 20:46
Gemini 3.6 Flash Is HERE – Testing Google’s BEST Model Yet!

Google's Gemini 3.6 Flash is a new model that offers significant speed improvements, a 17% reduction in token usage,...

Bijan Bowen summarized 2026-07-21 21:06
Kimi K3 Is INSANE – Is THIS a Sol & Fable Competitor?

Kimi K3 is a massive 2.8 trillion parameter open-weight model that claims to rival proprietary models like Claude...

Bijan Bowen summarized 2026-07-16 20:52
Meta Muse Spark 1.1 First Test – Is THIS a Frontier Model?

Meta's Muse Spark 1.1 shows benchmark improvements over its predecessor but fails in practical coding and game...

Bijan Bowen summarized 2026-07-11 16:41
I Tested GPT 5.6 Sol vs Fable (4 Real Uses Cases)

GPT 5.6 Soul outperforms Fable 5 in three out of four real-world coding and design tasks at a fraction of the cost,...

Jason Lee summarized 2026-07-11 10:32
GPT-5.6 Is HERE – Is THIS the Most INFURIATING Model Yet?

OpenAI released GPT-5.6 as a trio of models—Soul, Terra, and Luna—each with different cost and capability tiers,...

Bijan Bowen summarized 2026-07-10 01:43
Grok 4.5 explained in 8min..

Grok 4.5 is a 1.5-trillion-parameter frontier model from xAI, built on a new V9 foundation that incorporates...

Caleb Writes Code summarized 2026-07-10 01:04
Grok 4.5 Is INSANE – Is THIS a GPT & Opus Competitor?

Grok 4.5 is a highly competitive AI model that rivals state-of-the-art models like GPT-5 and Opus 4, offering strong...

Bijan Bowen summarized 2026-07-09 04:03
Field Guide to Fable — Thariq Shihipar, Anthropic

Anthropic's Thariq Shihipar introduces Fable, a new Claude model that represents a major leap in capability, akin to...

AI Engineer summarized 2026-07-06 16:00
Fable 5 is back..

Fable 5 is Anthropic's latest high-intelligence model, released after a 17-day suspension for safety overhauls, and...

Caleb Writes Code summarized 2026-07-03 00:52
Fable 5 is back… here is my plan

Fable 5 is back after a brief ban, and the speaker considers it a step-change in model capability, especially for...

David Ondrej summarized 2026-07-02 19:50
Why the Next 12 Months Will Create More AI Wealth Than the Last 100 Years

Three major AI companies (SpaceX, Anthropic, OpenAI) are filing for IPOs at a combined ~$4 trillion valuation,...

AI Founders summarized 2026-07-02 16:00
Claude Sonnet 5 Is HERE – Hands-On With Anthropic’s NEW Model!

Anthropic released Claude Sonnet 5 alongside a Linux desktop app beta. The model offers improved agentic coding and...

Bijan Bowen summarized 2026-06-30 23:29
GLM-5.2 vs Claude Opus 4.8 – Does GLM REALLY Beat Claude?

GLM 5.2, an open-weights model, is compared head-to-head against Claude Opus 4.8 across several challenging tasks: a...

Bijan Bowen summarized 2026-06-29 15:57
Ornith 1.0 First Look & Test – The BEST New Local Coding Models?

Ornith 1.0, a family of fine-tuned local coding models (9B dense and 35B MoE variants tested), shows promising...

Bijan Bowen summarized 2026-06-28 11:11
Introducing Ornith 1.0 - Agentic Coding LLMs

Ornith 1.0 from Deep Reinforce introduces self-scaffolding LLMs for agentic coding, where the model learns to...

Sam Witteveen summarized 2026-06-26 14:00
I Tested the Fable 5 Killer (Hermes Agent)

Fable 5 is not actually killed by Sakana Fugu or GLM 5.2; the tests show that Fugu is an intelligent router that...

Jack Roberts summarized 2026-06-24 16:59
North Mini Code First Test – Cohere’s LOCAL Agentic Coding Model!

Cohere's North Mini Code is a 30-billion parameter mixture-of-experts Apache 2.0 licensed coding model optimized for...

Bijan Bowen summarized 2026-06-20 13:20
OpenRouter Fusion First Test – Does THIS Beat Claude Fable?

OpenRouter Fusion, which synthesizes outputs from multiple models via a judge model, shows strong performance on...

Bijan Bowen summarized 2026-06-16 13:20
Omnigent: The New Meta-Harness for EVERY Coding Agent - Claude Code, Codex, Pi, More

OmniAgent is a new open-source meta harness from Databricks that orchestrates multiple AI coding assistants like...

Cole Medin summarized 2026-06-15 14:42
MYTHOS MYTHOS MYTHOS

Anthropic released Mythos 5 and Fable 5, a new class of 10-trillion-parameter models that exceed all previous models...

Matthew Berman summarized 2026-06-09 23:02
Anthropic Just Dropped Claude Mythos and Fable 5 (Full Breakdown)

Anthropic released Claude Fable 5 and Claude Mythos 5, with Fable 5 being a safer, public version of the previously...

Brock Mesarich | AI for Non Techies summarized 2026-06-09 19:35
Opus 4.8 Just Dropped. Here's How To Actually Use It.

Opus 4.8 is released with improved benchmarks and features, including sharper judgment, more honesty, and the...

Nate Herk summarized 2026-05-28 18:52
Finally a good benchmark (DeepSWE)

Deep Suite is a new benchmark for coding models that delivers four major advances over today's public benchmarks,...

Matthew Berman summarized 2026-05-27 16:03

Frontier News · by Hyperjump Technology