Blog
Articles on AI development, LLMO, context engineering, and harness engineering.
2026 113 posts
- Claude Code Max vs API vs Qwen: 4h Breakeven 18 min
- Gemini CLI vs Claude Code: Login Blocked Day 1 14 min
- nanochat GRPO Is Just REINFORCE + a 1-Line Regex 15 min
- 2.8x Traffic Spike: How to Identify Bot Traffic 22 min
- Bing Webmaster Tools vs Google: 787 queries 25 min
- Zenn noindex on Every Post: 8 Weeks Unnoticed 18 min
- Harness Engineering Benchmark: 13.7pt Same Model 10 min
- pgvector vs Qdrant vs Weaviate at 10k: No Cliff 16 min
- I Measured AI Citation Half-Life for 90 Days: 6 of 10 Sources Vanish by Week 4 21 min
- Static Call Graphs Miss 61% of Invocations 17 min
- AI Mode Cited My Portuguese 3.4× English. Japanese Also Beat English. 18 min
- Claude Skills vs Subagents in 2026: Which One To Reach For (7 Decision Rules) 18 min
- Claude Code vs ChatGPT Codex: 30-Day Cost by 7 Task Types 16 min
- Spec-Driven Development with Claude Code: 3 Ways the Spec Itself Broke Us 19 min
- Codex CLI vs Claude Code: 7 Real Tasks, Same Repo, 31 Days Later 17 min
- GitHub Copilot Agent Mode vs Claude Code: 8 Tasks, 31 Days, Real Bills 17 min
- Claude Code vs Aider on Same 10 Tasks — One Finished 8, the Other 5 14 min
- Claude Code vs Qwen3-35B on RTX 4070: 34.6 tok/s Break-even 17 min
- Karpathy's nanochat on a MacBook: End-to-End runcpu.sh Log from Tokenizer to Chat (2026) 13 min
- Claude Code Login: 2 Auth Layers vs setup-token 12 min
- The Shai-Hulud npm attack: why signature verification, npm audit, and --ignore-scripts all failed 8 min
- Can Claude Code Generate Diagrams? 3 Conditions 11 min
- Claude Code: The Operator's Field Guide 16 min
- llms.txt vs Claude Skills manifest: I tested 5 AI engines to see which file they actually read 18 min
- When Claude Code's Auto Mode Blocks Only Bash: Investigating the Safety Classifier Outage, Plus a Fail-Open-on-Outage Hook Design 16 min
- Chat Models Are Born in the Loss Mask — Reading nanochat's SFT 11 min
- Five Days of Tracking 346 Aftershocks From 90 km Away — and a Rematch With the 2016 Kumamoto Earthquake 7 min
- nanochat: $48 to Retrain GPT-2 (8×H100, 2026) 16 min
- JMA Earthquake Prediction Impossible: Day 8 Miss 23 min
- llmoframework Audit on 30 Dev Blogs: 4 of Top 5 Fail Same 3 Checks 20 min
- Cursor Composer vs Claude Code: 400k-Token Repo Benchmarked 11 min
- nanochat: I Retrained GPT-2 for $48 in 2026 12 min
- PinchTab Only Shoots One Viewport: Teaching a 9.4k-Star Browser Bridge to Capture Full Pages 10 min
- pixelshot Read One Tile of an 18,609px Wikipedia Page: the Lazy-Load Trap Under Visual RAG 14 min
- The Machine Accent Travels: AI Text Is Rhythmically Monotone in 70/70 Cells Across 3 Languages 7 min
- Article Schema Alone Didn't Make AI Recognize Me as the Author. The Entity Wiring That Did (in 4 JSON-LD Fields). 17 min
- OpenCut MCP: Drive a Video Editor from Claude Code (4 Traps) 13 min
- MCP Server Audit: 4 Layers, A-F Grade 24 min
- The Skill Eval Repo I Didn't Build: 107 SKILL.md Files, 6 Checks, 21 False Positives 15 min
- Ship a product, get a support button for free: an edge-injected overlay on Cloudflare Workers 13 min
- FUNDING.yml alone won't show a Sponsor button: notes from auditing 42 repos 8 min
- historymap: one YAML file becomes a corporate-style product-history timeline 10 min
- Claude Code Made Me 40% Slower: 3 Places 18 min
- The 7-Step Test That Told Me When to Switch From RAG to GraphRAG 18 min
- I Ran Claude Code, Cursor, and Codex Side by Side for 31 Days. Here Is the Real Monthly Bill. 20 min
- Claude Code vs Cursor: 6 Tasks, 1 Uninstalled 17 min
- Claude Code, Cursor, Codex: 340k Token Breakeven 16 min
- llama.cpp --cpu-moe on RTX 4070: Qwen 2.8x tok/s 16 min
- Your AI Code Review Burns 80% of the Context Window on Files It Never Needed 19 min
- I Made Claude Code Review Only the Blast Radius — Token Bill Dropped 8-49x 16 min
- Anthropic frontend-design skill: #F4F1EA Named 19 min
- Claude Code, Cursor, Codex: 412 Decisions/Day 19 min
- Measuring AI Citation Half-Life: A 90-Day Methodology With 3 Real Decay Curves 21 min
- Multi-Agent Decision Fatigue: I Counted 412 Micro-Choices a Day. The Harness Cut It to 38. 17 min
- I Added 3 Numbers to One Paragraph. Perplexity Started Citing It in 11 Days. 22 min
- AI Mode Just Hit 1 Billion Users, and Opened a Local-Business LLMO Market Most Engineers Are Ignoring 13 min
- I Wired My Pages Into Topic Hubs, Not a Flat List: AI Citations Consolidated Onto 4 of Them 14 min
- AI Search Is Under 1% of My Traffic and 12% of My Signups. That's the LLMO Case I Actually Use. 15 min
- GEO's +115% From Statistics Is Domain-Dependent: It Worked for My Tech Posts and Did Nothing for My How-To Pages 14 min
- AI Reads Your Chunks, Not Your Page: I Promoted 9 Sections from H3 to H2 and Watched Which Ones Got Quoted 18 min
- I Wired Claude Code to Real Hardware Over USB Serial. The MCP Tool Was the Easy Part. 18 min
- I Checked What GPTBot Actually Sees on My JS-Rendered Pages. It Was an Empty `<div>`. 13 min
- AI Wrote 100 Passing Tests. Mutation Testing Says They Caught 58% of Real Bugs. 12 min
- AI Finds Your Page Three Ways. I Published the Same Fact in All Three and Timed Which Reached AI First. 14 min
- Claude Code Pricing vs API vs Local GPU in 2026 14 min
- Multilingual LLMO: 4 Languages, Wrong Citation 12 min
- I Rewrote 12 Pages to Answer the Question in the First Sentence. AI Started Quoting 7 of Them. 13 min
- I Rank #1 on Google. On Brave I'm Page 5. My Own AI Agents Can't Find Me. 18 min
- Perplexity Citations Exploded After I Changed 3 Things. Only 1 Was Schema. 20 min
- I Gave Every Page on My Site a .md Twin. The AI Fetchers Stopped Guessing 13 min
- My Best Page Went Stale in a Month: Why AI Search Rewards Freshness, Not Just Schema 12 min
- AI Search Splits Your One Question Into Six. My Pages Answered None of Them. 14 min
- I Stopped Adding Context to My Agent and Pruned Tool Outputs Instead — My 3-Hour Task Stopped Forgetting Its Own Plan 13 min
- 9-Week LLM Citation Decay: 46% Left, Google Flat 18 min
- Blast Radius: The File That Breaks Is 2 Hops Out 14 min
- Link-less Brand Mentions Beat Backlinks for AI Visibility — I Read the Ahrefs 75,000-Brand Study So You Don't Have To 15 min
- Your Page Rank Is Invisible to AI — Only Your Passages Get Cited 15 min
- I Crosspost to 4 Platforms with rel=canonical Pointing Home. AI Search Still Picks the Copy. 16 min
- Claude Code Skills Cost Tokens Even When They Don't Fire. I Measured 5 Skills Across 7 Hours. The Bill Was 18%. 17 min
- I Cron-Scheduled 7 AI Agents. 2 Silently Failed for 18 Days. Tracing Wouldn't Have Caught It. 19 min
- I Ran 3 Claude Code Sessions in Parallel for 8 Hours. They Overwrote Each Other's Context Twice. 18 min
- I Asked 5 AI Search Engines to Cite My Own Blog. Only 3 of 31 Articles Showed Up. 16 min
- I Added 11 JSON-LD Schemas. Three Months Later, Only 3 Showed Up in AI Citations. 18 min
- I Refactored 100 Functions With Claude. 7 Got Slower in Production. 17 min
- Claude Code Hooks for TDD: 4 of 10 to 9 of 10 20 min
- I Added a 4th Agent That Audits My Other Agents. It Caught My Strategist Procrastinating for 3 Weeks. 23 min
- I Translated My Blog Into 4 Languages. Portuguese Got Nearly 4× the Traffic of English. 14 min
- TRM 8,337% LLMO Playbook: 1 of 4 Pillars Worked 19 min
- Claude Said 'You're Absolutely Right!' 47 Times Last Week. I Was Only Right 11 Times. Claude Was Wrong 36. 12 min
- I Plugged the Same Site Into 7 AI-Citation Trackers. They Reported 7 Different Numbers. 15 min
- The 5 AI Crawlers That Hit My Sites Most in 30 Days — What Their Logs Told Me About LLMO 15 min
- I Plugged Claude into a Chaos Engineering MCP Server. It Killed Staging 4 Times Before Finding a Bug We'd Missed for 6 Months. 31 min
- I Caught Claude Hiding My Bug 3 Times in a Row. Then I Turned 10 Debugging Habits Into Prompts. 18 min
- I Gave My Strategist Agent WebSearch. 5 Topics Took 20 Minutes. Splitting It Into 3 Made It 3. 15 min
- I Benchmarked 5 Voice AI Stacks. Only 2 Stayed Under 300ms. 16 min
- Claude Code Sub-Agents Disagreed on 40% of My PR 18 min
- llms.txt Audited: 30 Files, 5 Anti-Patterns 19 min
- OpenClaw Hit 250K Stars Faster Than React. I Spent a Day Switching From Claude Code 18 min
- I Refused to Write Specs Until Claude Code Generated Wrong Code Three Times 15 min
- I Let My Claude Code Agent Run for 24 Hours. The $400 Bill Was the Least Scary Part. 19 min
- The og:type Bug Three of My Astro Sites Quietly Shipped 11 min
- I Stacked 4 More Context Layers on Top of RAG. The Improvement Was 12%. 14 min
- arXiv 2603.25723 "Natural-Language Agent Harnesses": 4 Patterns + 3 Anti-Patterns After 12 Weeks in Prod 22 min
- Claude Code Skills vs Commands: 47 Prompts 17 min
- Your New Domain's First Week of GA4 Is a Lie: 4 Days of Raw Data from kaoriq.com's Launch 11 min
- ChatGPT Codex vs Claude Code: 47 PRs, $297 Bill, Aug 2026 26 min
- One Question, Five AI Search Engines, Five Different Answers 19 min
- Is AI Actually Citing Your Site? How to Measure What Google Rankings Can't 16 min
- Generative Engine Optimization: 3 of 9 Worked 18 min
- 9 Bugs in My AI Pipeline: None Were the AI's Fault 9 min
- Claude Haiku + RAG Beat Sonnet 11.8 to 5.3 12 min
- llms.txt: The File That Decides Whether AI Can Find Your Site 13 min
- Building an Autonomous Content Pipeline with Claude Code 2 min