Work

I run an AI consulting firm and I build things with my hands. This page holds the work I can show, and every claim on it is checkable: the code runs, the links are live, and the numbers come from scripts you can download and rerun.

Client systems ship under NDA, so they are not listed here. The patterns from that work are public in the essays.

alexander@aros:~$ ls work/
01

AROS, the portfolio that is its own proof

The site you are on is the flagship project. Instead of a page that claims frontend skill, it is a desktop operating system built from scratch in vanilla HTML, CSS, and JavaScript. Anyone can view source and check the claim.

0lines hand-written
0npm dependencies
0windows
0skills tracked live

The build

A window manager with focus, z-order, drag, tile, and cascade. A boot sequence that plays once and gets out of your way. A terminal that answers more than 25 commands, including stacktop, a live skills monitor rendering 44 technologies with tier badges, load bars, and a grep filter. A ⌘K command palette. A release-notes system that greets returning visitors with what shipped since their last visit. There is no framework and no build step: package.json lists zero dependencies, and the whole thing is served by a 58-line Node static server on Azure App Service with its own MIME map, cache policy, and www-to-apex redirects.

The part that fought back

The desktop has a browser window showing the real ryshe.com homepage. Naive full-page screenshots ghost fixed elements and break 100vh layouts, so the capture pipeline drives headless Chrome through real scroll positions, grabs overlapping viewport segments, and stitches them with Pillow while cropping the sticky header out of every segment after the first. Small problem, real engineering.

JavaScriptCSSHTMLNode.jsAzure App ServicePlaywrightPillow
02

The LLM Tools Suite

Five free, no-signup tools for people shipping LLM systems. Each one exists because I kept rebuilding it for client work. Everything runs in your browser: no backend, no accounts, nothing logged.

0tools
0server calls
0models priced
0providers tracked

What they do

The MCP Server Generator turns a form into a working Python FastMCP or TypeScript server plus the Claude client config. The Context Budget Calculator and API Cost Calculator share one hand-verified price dataset covering 11 models across Anthropic, OpenAI, and Google, with a dated changelog on the model reference page. The Chunking Visualizer shows exactly where each splitter cuts and how much overlap duplicates. The Judge Rubric Builder produces checkable eval criteria, a judge prompt, and a JSON schema.

The engineering bit

Every tool serializes its full state into a base64url fragment, so a configured tool travels as a link and restores from localStorage on return. Shared pricing lives in one dataset file that three pages render, so a price change is one edit and one cache-buster bump.

Vanilla JSFastMCPMCP TypeScript SDKzodlocalStorage
03

The chunking benchmark: folklore, measured

Most chunking advice is vibes. I measured it: three splitting strategies on 20 SQuAD articles and 800 real questions, embedded with all-MiniLM-L6-v2, scored as answer recall at k, on a laptop CPU. The script is published so anyone can rerun the whole thing.

0questions
0strategies
0index tokens, identical across strategies
0recall@1, recursive over fixed

Headline numbers, no overlap

Strategyrecall@1recall@5Index tokens
Fixed-size61.1%87.4%161,426
Sentence66.9%88.9%161,426
Recursive70.8%89.6%161,426

The surprise

Overlap helps recall but overshoots badly at span granularity: asking for 30 tokens of overlap produced 17.5% duplication on fixed chunks, 44.6% on sentence chunks, and 78.3% on recursive chunks, because overlap gets rounded up to whole sentences and paragraphs. And by k=5 the three strategies nearly converge, which means rerank depth can rescue a bad splitter.

Pythonsentence-transformersall-MiniLM-L6-v2SQuAD v1.1
04

The content engine

Thirteen essays in, nothing here publishes when inspiration strikes. The writing runs like an engineered system: topics chosen from real search demand, recurring series with deadlines, and a pipeline where an essay ships with its card, schema, feed entry, and release note in one motion.

0essays, all cluster-mapped
0series on deadlines
0original dataset published
0auto-generated posts

The system

Every essay starts from a question engineers actually search for, then gets mapped to a cluster so the pieces reinforce each other instead of competing. Two series run on deadlines: Failure Modes, one production failure mode a month on first Tuesdays, and a quarterly data report built on an original benchmark. One shared manifest file drives the homepage cards, the essay index, next-essay navigation, and the About window's count, so publishing is a single ritual instead of five chances to forget something. Each essay ships with a script-generated social card, FAQ schema where it earns one, an RSS entry, and a line in the release notes.

The rule

It is written for engineers who can smell filler: no auto-generated posts, no keyword mush. And when an essay makes a quantitative claim, the script that produced the number gets published next to it.

Technical SEOschema.orgPillowRSS
05

The ship loop: AI-assisted, human-gated

This site ships almost daily, with an AI agent doing much of the labor and a set of hard guardrails deciding what it is allowed to do. That loop is what I help clients build at Ryshe, and here it runs in production on my own property.

0AI co-authored commits
0agent pushes, ever
100%deploys hash-stamped
0benchmark script published

The guardrails

The agent writes code, runs builds, and deploys, but it cannot push to the repository: a hook refuses the command, so a human runs every push. Fabricated numbers are banned, which is why the chunking data on this page came from a benchmark I ran rather than a plausible-sounding guess, with the script published alongside the report. Every deploy is stamped with its git hash, which the site footer reads live from /version.txt, and every meaningful change lands in the release notes.

Why it is on this page

44 of the 45 commits behind this site carry an AI co-author trailer, and every one of them was reviewed before it landed. The interesting engineering is not the prompt. It is the system around the model: what it may touch, what requires a human, and how claims get verified before they go live. That is the same discipline I bring to client AI adoption.

Claude Codegit hooksAzure CLIRelease discipline
06

Ryshe, the company and its platform

Ryshe is my AI and cloud consulting firm. It exists to bridge the gap between AI promises and IT reality: strategy, retrieval systems, evaluation pipelines, MLOps, and cloud architecture, delivered with honest engineering. I run it, and I also built its platform end to end.

0linked domains
0publishing pipeline
0manual rebuilds
100%static delivery

The platform

ryshe.com is an Astro static build on Azure Static Web Apps, with DNS on Azure. Publishing is webhook-driven: when content lands, the site rebuilds itself, so nobody at Ryshe runs a deploy by hand. The architecture is deliberately split across two properties: ryshe.com speaks to buyers, this site speaks to engineers, and structured data links the two so search engines understand they share an author.

The practice

Client systems ship under NDA. What I can show is the thinking: the essays here document the patterns from that work, from evaluation pipelines to observability for silent degradation, and the tools are the utilities I kept rebuilding on engagements, made public.

AstroAzure Static Web AppsAzure DNSPythonTypeScript

The full technology list lives in the Tech Stack window on the desktop, or type projects in its terminal. For what shipped recently, see the release notes.