← Zalman Friedman

How This Site Works

Problem
A job-search portfolio needs to work for a recruiter skimming for 30 seconds, and also for a hiring manager digging for 20 minutes. Plus, I wanted it to demonstrate my AI-forward product judgment.
My role
All of it: product decisions, content, design, and the build (working AI-natively).
Outcome
A plain-HTML site with an AI assistant grounded in one document, costing under $5/month with a hard spend ceiling.

The design decision: tasteful but restrained

This site is static HTML with one stylesheet. No framework, no build pipeline, and no trackers. This speaks to the same judgment I apply to products: complexity is a tax you should only pay for a solid return. A portfolio's job is to deliver information fast; nothing delivers information faster than text on a page.

The AI decision

The assistant on the homepage answers questions from a single career document. While RAG might be a good choice for corpora too large to fit in a model's context window, my career document is only about 25,000 tokens. The entire document ships with every request as a cached system prompt.

That ensures we have no retrieval misses (the model always sees everything), no extra infrastructure to break, and deterministic behavior. Prompt caching makes it cheap: after the first request, the document costs about a tenth of the normal input rate. Choosing the boring architecture over the impressive-sounding one is the grounded product judgment so elusive in today's landscape.

The cost model

The assistant runs on Claude Haiku. Per message: roughly 15K cached input tokens plus a 400-token capped reply comes to about a third of a cent. At realistic job-search traffic that's a couple of dollars a month. But "realistic" is only a hope, so the actual controls are layered:

Questions are capped at 400 characters and replies at 400 tokens. Each visitor gets 10 questions per session and 25 per day (tracked per IP). A global kill switch stops the assistant after 500 messages in a day and falls back to pointing at the resume. Above all of that, a hard monthly spend limit at the API account level means the absolute worst case is an amount I chose.

The security posture

Someone will try to jailbreak this bot, maybe today. The defense isn't a clever prompt, but that there's nothing behind the boundary: the bot's document is the public, recruiter-safe version of my career file, containing nothing I wouldn't say in an interview. A fully successful jailbreak yields an off-topic answer at my expense of a third of a cent. The system prompt still instructs the model to stay on topic, refuse persona changes, and decline salary questions, but the real control is that the worst case is contained.

This is the same principle I applied scoping the MCP server at Propela: don't try to make the agent unbreakable; decide what it's allowed to touch so that breaking it doesn't matter.

The stack, for completeness

Cloudflare Pages (static hosting, free) · one Pages Function for the chat endpoint · Cloudflare KV for rate-limit counters · Cloudflare Turnstile for bot filtering · the Anthropic API (Claude Haiku) with prompt caching. Total infrastructure cost: the domain, plus single-digit dollars a month for inference, capped.