Hire a senior AI engineer who has shipped production features against an LLM bill.

Not a prompt copy-paster, not an AutoGPT enthusiast. A real engineer with twelve years of practice and live production AI features running on this site right now: the LLM citation tracker, the AI search tools, the content brief generator, the entity-schema layer. Direct engagement, fixed price, no agency markup.

12 years senior engineering Production AI features live on this site Anthropic Claude API in production since 2024 Direct engagement, no marketplace markup

KEY FACTS · 2026

  • An AI engineer in 2026 is a real software engineer who builds production features against LLM APIs (Anthropic Claude, OpenAI, Google Vertex) and agentic frameworks. Not a prompt copy-paster, not an AutoGPT enthusiast: an engineer who ships RAG, agents, tool use, and prompt caching into production codebases.
  • Hourly rates in 2026: 150-300 USD/hour for senior independents who have shipped real AI features into production. Mid-level: 90-180 USD/hour. Anyone offering AI engineer rates under 60 USD/hour is reselling overseas labour around a model wrapper. The production failure rate matches.
  • Five production-grade skills the role requires: (1) prompt caching + token-budget discipline, (2) tool use + structured output for agentic workflows, (3) RAG with proper chunking + reranking, (4) eval frameworks and regression tracking, (5) cost monitoring on a real Anthropic / OpenAI bill.
  • Twelve years of senior engineering practice. Anthropic Claude API in production on this site since 2024 (free public tools, content pipelines, agentic admin actions, the LLM citation tracker that runs every Monday). The Claude Code build pattern documented end-to-end.
  • Direct-hire skips agency overhead. Engagement shapes available: feature build (1-2 weeks, 5-15k USD), audit + architecture (3-5 days, 3-8k USD), sprint (4-8 weeks, 15-45k USD), platform (12-20 weeks, 45-150k USD).

SKILLSET, WHAT YOU'RE ACTUALLY HIRING

Anthropic Claude API integration

Sonnet 4.6, Opus 4.7, Haiku 4.5 in production. Prompt caching to cut bills 75% on repeat prompts. Tool use for agentic features. Files API. Citations. Batch API for offline processing. Vision for image-input features.

Agent SDK + agentic patterns

Anthropic Agent SDK in production. Sub-agent delegation, hooks for production safety, skills + memory for long-running sessions. Plus the supporting infra: rate limiting, cost monitoring, error replay, output validation.

RAG (retrieval-augmented generation)

Vector database choice (pgvector, Pinecone, Qdrant) by use case. Document chunking strategy. Hybrid search (BM25 + dense). Reranking with cross-encoders. Evaluation: retrieval precision, answer accuracy, citation correctness. Built more than three production RAG systems.

OpenAI + Google Vertex integration

When the use case fits, GPT-5 / 4o or Gemini 2.5 Pro / Flash. Multi-provider patterns: response caching across providers, fallback on rate limits, comparable eval harnesses. Not every project needs Claude; the engineer picks the model.

Eval frameworks

No shipping AI features without an eval set. Promptfoo, Inspect AI, or hand-rolled. Regression tracking on model upgrades. The smoke-test pattern that catches Sonnet → Opus drift before it reaches production.

Next.js + Supabase + Vercel

The default production stack for AI features in 2026. Edge functions for streaming. Vercel AI SDK for chat UIs. Supabase Postgres + pgvector for RAG. Realtime for collaborative AI features. Stripe for billing AI-feature usage.

Claude Code daily operation

Building AI features with AI tooling. Skills, sub-agents, hooks, MCP servers, memory. The same agentic patterns I ship for clients I use to build for clients. This site is the working example.

Cost monitoring + production safety

Real Anthropic / OpenAI bills monitored. Cost-per-feature broken out by use case. Hard limits per IP / per user / per organisation. Rate limiting that survives a traffic spike. Hooks that refuse to call paid endpoints without cache headers.

ENGAGEMENT SHAPES + 2026 PRICE BANDS

PROOF, PRODUCTION AI FEATURES LIVE ON THIS SITE

The fastest way to verify an AI engineer is to look at the AI features they have already shipped. The full list is built and operating on this site right now.

  • LLM citation tracker, weekly cron pulls real ChatGPT, Perplexity, Gemini responses via DataForSEO, parses citations, scores target-domain visibility. Production cost: ~$10/month.
  • AI Citation Checker, free public tool. Derives 5 buyer queries with Claude, runs each through Google with AI Overview detection. ~$0.03/run.
  • AI Search Keyword Research, harvests PAA + Reddit + Stack Exchange, clusters with Claude, returns AI Overview presence + Perplexity answer with sources.
  • AI Content Brief Generator, DataForSEO SERP + Claude synthesis → structured Markdown brief with common themes, required entities, H2/H3 outline. ~$0.05/run.
  • Business Name Generator + Trademark Check, Claude generates 20 candidates, Signa.so trademark search across USPTO + EUIPO + WIPO Madrid on top 12.
  • The entity-schema layer on every public page, Claude-translated 30-language content with humaniser pass, on the sister Deluxe Astrology property (91,000+ pages).

FREQUENTLY ASKED QUESTIONS

What is an AI engineer in 2026?

A real software engineer who builds production features against LLM APIs and agentic frameworks. They ship prompt-cached endpoints, RAG systems, tool-using agents, eval harnesses, and cost-monitored AI features. They are not prompt copy-pasters; they are not AutoGPT enthusiasts. The distinguishing trait is having a live production AI feature with a real bill against an LLM provider.

How is an AI engineer different from a Claude Code developer?

Overlapping but distinct. A Claude Code developer uses Anthropic Claude Code as their primary execution tool to build any software (whether or not the output uses AI). An AI engineer builds production software where AI is part of the user-facing feature. The same person often does both, I do, but the engagement shapes are different. Claude Code work scales horizontally (more code, faster). AI engineer work scales vertically (smarter features, lower per-task cost).

What is the hourly rate for an AI engineer in 2026?

Three bands. Senior independents who have shipped real production AI features: 150-300 USD/hour. Mid-level engineers with 1-2 years of AI work: 90-180 USD/hour. Anyone advertising rates under 60 USD/hour for "AI engineering" is reselling overseas labour wrapping an OpenAI key. Production failure rate matches the rate.

What does an AI engineer build?

Five common project shapes. AI search and citation tools (the AI Citation Checker on this site is one). RAG-backed Q&A over a documentation corpus. Agentic admin actions (a button that delegates a multi-step task to Claude). AI-augmented content pipelines (DataForSEO + Claude → published draft via human review). Customer-facing AI features (smart classifiers, recommendation engines, semantic search). The skill set transfers across all five.

How long does AI feature development take?

Most production AI features ship in 1-2 weeks for a focused build, 4-8 weeks for a multi-feature sprint, 12-20 weeks for a platform. The unknown is not the AI part (that ships fast with current APIs) but the eval set, the production-safety harness, and the cost-monitoring infrastructure. Anyone promising a 1-week timeline for a "production-grade RAG system" is skipping the eval set, the rate limiting, or both.

Can you hire an AI engineer directly without going through an agency?

Yes, direct engagement is the default here. Three shapes. Solo senior consulting: audits, architecture reviews, fixed-price feature builds where I am the named engineer. Agency engagement through Seahawk Media: for larger builds where you want a team behind the senior, with project management and QA. Hourly retainer: for ongoing capacity with no defined endpoint. Disclosure: I co-founded Seahawk so agency-routed projects include the standard agency cost structure; direct skips that overhead.

What stacks do you build AI features in?

Primary stack. Next.js 15 + React 19 + Server Components on Vercel. Supabase Postgres + pgvector for storage and RAG. Vercel AI SDK for streaming chat. Anthropic Claude (Sonnet 4.6 for bulk, Opus 4.7 for flagship, Haiku 4.5 for cheap classification). OpenAI GPT-5 / 4o when the use case fits Claude badly. Promptfoo or Inspect AI for evals. Stripe for billing AI usage. Resend for transactional email triggered by AI flows.

What about agents, is that real or hype?

Real in 2026, hype in 2024. Agents work in production when three conditions hold. First, the task is bounded (book a meeting, summarise five URLs, run an audit) rather than open-ended (build me a startup). Second, the eval set covers failure modes. Third, the agent has hooks that gate destructive actions. The Claude Code Agent SDK shipped in 2025 makes this tractable. Anthropic Claude is currently the strongest agentic model; OpenAI is closing the gap.

Will my AI feature still work in 6 months?

Yes if you build evals on day one. Model upgrades happen every 3-6 months. Drift is real (Sonnet 3.5 → Sonnet 4 → Sonnet 4.6 each shifted output style). An eval harness with 50-100 cases catches drift before production. Without one, you discover the regression from customer complaints. Production AI features without eval coverage are a ticking clock; the engineer who builds the feature should build the eval set in the same engagement.

Why hire you specifically?

Three reasons. Twelve years of senior engineering practice, AI is layered on top of real production experience, not a substitute for it. Working production AI features visible on this site right now: the LLM citation tracker, the AI citation checker, the AI search keyword research tool, the AI content brief generator, the entity-schema layer, the headless WordPress agency picker. And I built all of them with Claude Code, so I can debug both the AI feature and the AI tooling that produced it. If you want an AI engineer who ships the work he talks about, this is it.

WHAT THE FIRST 48 HOURS LOOK LIKE

Book a 30-minute call. Bring the AI feature you want to ship (or the problem you want to solve, if the AI part is not yet defined), your stack, and the deadline. By the end of the call you will know whether the project fits direct-hire or agency engagement, what the eval + cost-monitoring harness looks like for your specific use case, and a fixed price for the work. If AI is the wrong tool for the job I will tell you that.

RELATED