Archive
-
2026-10-02
:: pi-delegate: Claude does the thinking, a cheaper model does the typing
::
#agentic-ai,
#ai,
#llm
-
2026-09-25
:: A year of Fiona: my personal assistant, the Hermes swap, and the agent that rewrote its own playbook
::
#agentic-ai,
#ai,
#llm,
#self-hosted
-
2026-09-21
:: MTP on Strix Halo: +79% at one stream, zero at four, and the ceiling isn't bandwidth after all
::
#ai,
#hardware,
#llm,
#self-hosted
-
2026-09-21
:: A Strix Halo engine audit: a scary claim that didn't reproduce, a fork worth +35-92% prefill, and seven broken routes
::
#ai,
#hardware,
#llm,
#self-hosted
-
2026-09-20
:: Would Jev help pi-rukas? Mostly no, and the reason is the interesting part
::
#agentic-ai,
#ai,
#llm
-
2026-09-17
:: DwarfStar, DiffusionGemma, and the ROCm 10 unfreeze: an ecosystem status report from halo
::
#ai,
#llm,
#self-hosted
-
2026-09-17
:: Four coding models on one Strix Halo: pass@1 hides two very different failure modes
::
#agentic-ai,
#ai,
#hardware,
#llm,
#self-hosted
-
2026-09-17
:: MTP on halo: two production adoptions, one banned backend, one corrected claim
::
#ai,
#hardware,
#llm,
#self-hosted
-
2026-09-17
:: Do synthetic LLM benchmarks lie? On halo, real code matched llama-bench to 0.6%
::
#ai,
#hardware,
#llm,
#self-hosted
-
2026-09-16
:: Cerebras vs our H100: 11x the decode speed, half the agent turns
::
#agentic-ai,
#ai,
#llm,
#self-hosted
-
2026-09-15
:: An afternoon of agentic coding: 12,53 € on our H100, 47 € to 726 € through an API
::
#agentic-ai,
#ai,
#llm,
#self-hosted
-
2026-09-15
:: LFM2.5-2.6B, the silent submarine, and why retrieval doesn't judge what it retrieves
::
#agentic-ai,
#ai,
#llm,
#self-hosted
-
2026-08-31
:: Qwen3.8-27B-INT4 with MTP: 30-40% faster round trips, and two rival explanations for the divergence
::
#agentic-ai,
#ai,
#hardware,
#llm,
#self-hosted
-
2026-08-30
:: Qwen3.8-Flash-Next on Strix Halo, part 2: the missing speed was already merged upstream
::
#ai,
#hardware,
#llm,
#self-hosted
-
2026-08-28
:: Qwen3.6-35B-A3B on Strix Halo: one flag buys 22% prefill, and my own GPU power advice was wrong
::
#ai,
#hardware,
#llm,
#self-hosted
-
2026-08-28
:: strix-halo-optimize: a tuning skill whose rules had to become code
::
#agentic-ai,
#ai,
#hardware,
#llm,
#self-hosted
-
2026-08-27
:: Qwen3.8-Flash-Next on Strix Halo: a 177B Qwen4 preview at 22 tok/s, one day after release
::
#ai,
#hardware,
#llm,
#self-hosted
-
2026-08-21
:: MTP speculative decoding costs code correctness on two Qwen-lineage models
::
#agentic-ai,
#ai,
#llm,
#self-hosted
-
2026-08-21
:: Ornith 1.5 35B A3B on Strix Halo: MTP pays off at n=1, not at the default n=3
::
#agentic-ai,
#ai,
#hardware,
#llm,
#self-hosted
-
2026-08-12
:: The small model didn't fabricate. It stopped citing.
::
#agentic-ai,
#ai,
#llm,
#self-hosted
-
2026-08-07
:: AMD buys Taalas: etched models, and what it means for local inference
::
#agentic-ai,
#ai,
#hardware,
#llm,
#self-hosted
-
2026-07-27
:: Poolside says Laguna needs an H200. We ran it on a single H100.
::
#agentic-ai,
#ai,
#hardware,
#llm,
#self-hosted
-
2026-07-26
:: DeepSeek-V4-Flash on Strix Halo: it runs, and now we know how fast
::
#ai,
#hardware,
#llm,
#self-hosted
-
2026-07-26
:: Laguna-S-2.1 on a mini-PC: the honest numbers
::
#agentic-ai,
#ai,
#hardware,
#llm,
#self-hosted
-
2026-06-18
:: Loop engineering: the term is two weeks old, the practice is over a year old
::
#agentic-ai,
#ai,
#multi-agent,
#tools
-
2026-06-18
:: Adding an agent role is more expensive than it looks
::
#agentic-ai,
#ai,
#multi-agent,
#research
-
2026-06-15
:: Running an H100 at Trail Openers: what it actually costs in money, energy, and CO₂
::
#agentic-ai,
#ai,
#hardware,
#llm,
#self-hosted,
#sustainability
-
2026-06-05
:: H100 vs Strix Halo: the gap is bigger than the first benchmark suggested
::
#agentic-ai,
#ai,
#hardware,
#llm,
#self-hosted,
#sustainability
-
2026-05-21
:: Live steering breaks deep focus: notes from three failed pair-coding sessions
::
#agentic-ai,
#ai,
#multi-agent,
#tools
-
2026-05-20
:: pi-ensemble: agentic coding that optimizes for quality, not speed
::
#agentic-ai,
#ai,
#open-source,
#tools
-
2026-05-20
:: Skills, not features: notes on a methodology that works for one person at a time
::
#agentic-ai,
#ai,
#methodology,
#open-source
-
2026-05-13
:: vipune 0.5: agent memory without the agent
::
#agentic-ai,
#ai,
#rust,
#tools
-
2026-04-30
:: New Strix Halo? Five things that will cost you hours.
::
#ai,
#hardware,
#llm,
#self-hosted
-
2026-04-28
:: You feeling the squeeze yet?
::
#agentic-ai,
#ai,
#cost
-
2026-04-27
:: Feynman: a research agent worth the rough edges
::
#ai,
#research,
#tools
-
2026-04-27
:: System prompts steer. Permissions stop.
::
#agentic-ai,
#ai,
#security
-
2026-04-26
:: BYOT: Bring Your Own Tokens
::
#ai,
#compliance,
#security
-
2026-04-26
:: Hello, world.
::
#meta