Skip to content

Blog

Notes on AI, custom agents, and building software that works. The long-form version of what I share on X and LinkedIn.

7 min read

Your VM Is Under SSH Brute-Force Right Now: My Linux Hardening Baseline

An empty VM took 10,304 failed SSH logins in 24 hours. Key-only SSH verified in the log, a provider firewall Docker cannot bypass, fail2ban, secrets, CI keys, and a laptop that is part of the server.

Read post
3 min read

ClickFix: The Fake Cloudflare Check That Asks You to Press Win+R

A hacked WordPress site showed my mom a convincing Cloudflare lookalike asking her to press Win+R and paste a PowerShell command. How ClickFix works, and the one rule that protects your family.

Read post
4 min read

Why Public LLM Benchmarks Don't Match the Model You Run Locally

Published scores are measured on full-precision weights on server GPUs; your local 4-bit quant is a different model. The air-gapped project, the 70GB ceiling, and the five Apache 2.0 finalists.

Read post
4 min read

MLX, GGUF, MXFP4, NVFP4: Making Sense of Local LLM Formats

A 32GB RTX PRO 4500 Blackwell at the client reframed the benchmark. Model, precision, packaging and engine: why MLX versus GGUF is a crooked comparison and FP4 acceleration is not automatic.

Read post
6 min read

When a Local LLM Cites Documents It Never Opened: Two Benchmark Failure Modes

Granite 4.1 30B wrote the most authoritative answers of the round and fabricated citations in 12 of 20. Nemotron answered clean and fast by barely reading. The harness that caught both.

Read post
4 min read

Local Speaker Diarization with Whisper and pyannote: My Full Client Meeting Flow

A 41 minute client call became a speaker separated transcript in 87 seconds, all on my Mac: whisper-large-v3-turbo, the ivrit.ai Hebrew finetune, and pyannote.

Read post
5 min read

Detect Missed Calls Without READ_CALL_LOG: What 3 Play Store Rejections Taught Me

Three identical rejections, a denied appeal, then a one hour rebuild on NotificationListenerService that passed review. Now one instruction line forces the route comparison before any code.

Read post
4 min read

Claude Scrambles Hebrew Text? One Instruction Fixes Mixed RTL Lines

When a Hebrew line opens with an English word, Unicode BiDi flips the whole line. One user instruction stops Claude from ever triggering it.

Read post
5 min read

Video vs Image Tokens in Gemini: The Same Frame Costs 70 or 1120

The same frame costs 70 tokens inside a video and 1120 as an image. What media_resolution and fps actually control, and when to send frames instead of video.

Read post
3 min read

I Stopped Taking Notes in Client Meetings: Local Transcription with mlx-whisper

I recorded a 44 minute client kickoff instead of taking notes and transcribed it on my own Mac with mlx-whisper and the ivrit.ai Hebrew model.

Read post
5 min read

Model Judgment Is Not Access Control: One Sentence Broke My Agent

My agent read the terms, refused for the right reason, and then complied once I claimed a permission I never had. Constraints belong under the model.

Read post
4 min read

Stop Accidental Claude Code Approvals: CLAUDE_CODE_DISABLE_MOUSE_CLICKS

Clicking the terminal to focus it can land on a prompt answer and approve something you never wanted. One environment variable fixes it.

Read post
5 min read

How to Give Claude Code a Project Memory: The Devlog Plugin

My tickets, decision log, and progress log method is now devlog: a free, MIT-licensed Claude Code plugin. No hooks, no scripts, just markdown.

Read post
6 min read

How I Built a Free PDF Editor in 15 Minutes with Claude Code

My wife was paying $12 a month for a shady PDF site. Fifteen minutes with Claude Code later: a free, private PDF editor that runs entirely in the browser.

Read post
3 min read

How to Stop Claude Code From Forcing Git Worktrees

Background sessions are forced into a fresh worktree with no .env and no node_modules. One setting, worktree.bgIsolation, turns that off.

Read post
5 min read

A CI Dead Man's Switch for Tech Debt: The Allowlist That Can't Rot

A lint allowlist that fails CI both ways: new violations break the build, and so does fixing the code without clearing the marker.

Read post
5 min read

Markdown Tickets in the Repo: Issue Tracking for AI-Agent Projects

No Jira, no GitHub Issues: tickets are markdown files where the filename is the ID, and a Paid field in the header decides who fixes them.

Read post
5 min read

progress.md: How to Know What Your AI Coding Agent Actually Built

A live status mirror of the spec with one rule: nothing is done without proof. Four statuses, real evidence, failures included.

Read post
6 min read

Why Your AI Coding Agent Needs a decisions.md File

A single append-only file at the project root gives Claude Code the one thing git can't: why you decided. The setup, a real entry, and its limits.

Read post
3 min read

Production LLM Tip: Run the Weakest Model That Still Does the Job

It sounds backwards, but I work hard to run production on the weakest model that does the job. It is cheaper, and it leaves a reserve for the day things break.

Read post
2 min read

Claude Filling Chrome with Tab Groups? Switch to agent-browser

Every session of the Chrome extension leaves another colored tab group behind. The fix is agent-browser: a separate Chrome, its own profile, its own accounts.

Read post
3 min read

Models Derby: 9 AI Models Compete on World Cup Predictions

9 models head to head on World Cup predictions: Claude, GPT, Grok, and two local models running on my Mac. Same prompt, hidden picks, one scoring table.

Read post
3 min read

My AI Development Setup: 4 Computers, One Giant Screen, Zero Regrets

Someone asked for the full setup breakdown. So here is everything: 4 computers, one giant screen, and the cheapest machine in the room doing the most work.

Read post
3 min read

remoteControlAtStartup: Pick Up Any Claude Code Session from Your Phone

With one setting, every claude you launch shows up in the mobile app with a green dot. I approved a refactor from the gym and came back to finished work.

Read post
3 min read

Claude Code Deletes Your Session History After 30 Days: The Fix

A default setting quietly deletes your local session transcripts after 30 days. No warning, no recovery. You may have already lost things. Here is the fix.

Read post
4 min read

Why Claude Code Quietly Stops Using Your Skills (and How to Fix It)

If you have many skills installed, Claude Code may have already stopped using some of them without telling you. Not a bug: a well-hidden default setting.

Read post
3 min read

How to Stop Claude Code Terminal Flicker with /tui fullscreen

Every streamed answer made the terminal shake. One hidden setting fixes it, keeps a 6-hour session at constant RAM, and makes the mouse work. One line to try it.

Read post
2 min read

The Eye: Monitoring Parallel Coding Agents from the macOS Menu Bar

Running several agents in parallel, I became their babysitter: jumping between windows to see who finished, who is stuck, who is waiting. So I built The Eye.

Read post
3 min read

Why an Open Model Insists It Is Claude: Distillation Fingerprints

You can learn a lot about how a model was built just by asking who it is. GLM-5.2 was sure it was Claude, and that is not a bug. It is a kind of paternity test.

Read post
2 min read

One Sentence, No Design Brief: Letting Claude Lead an App Redesign

Design is what sells, and I always neglect it. This time I gave Claude one sentence and let it run the process: questions, mockups, refinements, even dark mode.

Read post
3 min read

Surviving an Extension That Keeps Rewriting Your Claude Code Config

I wanted my custom status bar and the extension that pays you for showing ads. The extension kept winning. The full rabbit hole, and the fix that finally held.

Read post
2 min read

Is Your Website Readable by AI Agents? Mine Scored 34 Percent

An audit skill scored my site 25 out of 73 for agent readability. After one session with Claude it reached 78 percent. Agents are a growing share of your traffic.

Read post
3 min read

Where Local LLMs Actually Earn Their Keep: Three Real Agents

After the coding test flopped, I looked for where the local model does bring value. Three real agents from the last few days, and a rule for splitting work.

Read post
4 min read

Same Local Model, Opposite Settings: Configuring an LLM Per Task

Choosing the right model is half the job. The same Gemma 4 needs near-opposite settings for coding versus scanning data, and max context is usually a mistake.

Read post
2 min read

How to Connect a Coding Agent to a Local Model in LM Studio

A few people asked how to wire a coding agent to LM Studio. It is genuinely simple: load the model, copy the URL, add a custom endpoint in VS Code.

Read post
3 min read

Can a Local LLM Handle Real Coding? Testing Gemma 4 on My Mac

Plan and build a Snake game, on a model running on my laptop. It planned nicely, opened a sub-agent, then hit a wall a cloud model cleared in seconds.

Read post
3 min read

OPERATING_RULES.md: Run Your AI Agent in Auto Mode Without Losing Control

Approving every action breaks your finger, but full auto mode burns 300K tokens in half an hour. The fix is a rules file the agent must obey. My real rules inside.

Read post
2 min read

Adding Codex as a Second AI Code Reviewer in My Claude Code Workflow

When Claude and I are supposedly satisfied with a stage, Codex reviews the changes. Every point it raises goes back into the loop until it has no notes left.

Read post
2 min read

You Cannot Do Due Diligence on Software: A Fundraising Story

We declined the investor on Tuesday. On Wednesday we found a bug that made everything we had shown him look fabricated. A story about software and luck.

Read post
3 min read

Plan in Claude Chat, Execute in Claude Code: A Two-Layer AI Workflow

A Claude project holds the contract and scope, produces a build spec, writes each stage prompt, and reviews every Claude Code summary. Quality jumped.

Read post
2 min read

Enterprise AI Assistants: All-Knowing Copilot or Isolated Specialists?

One assistant that sees your email, Slack, and Drive, or a sandboxed specialist locked to the codebase. Least privilege is coming to AI assistants.

Read post
2 min read

A /document-project Command That Keeps AI Coding Agents Oriented

AI agents have no memory between sessions, so they guess. A slash command that re-documents the whole repo while you take a coffee break.

Read post

Don't miss a post

Subscribe with RSS, or follow me where the conversation happens.