I'm Aria Vance, a prompt engineer and LLM systems designer. I turn fuzzy ideas into reliable, production-grade AI behavior — across chat, agents, and automated pipelines.
I specialize in prompt architecture, evaluation design, and agentic workflows for teams building serious products on top of large language models. My work sits at the intersection of linguistics, systems thinking, and product judgment.
Before prompt engineering had a name, I was writing rule-based dialogue trees and fine-tuning small transformers. Today I help teams get more out of frontier models — GPT, Claude, Gemini, Llama — without a single line of retraining.
I care most about reliability under distribution shift: prompts that don't just demo well, but hold up against messy real-world input, adversarial users, and edge cases nobody thought to test.
Prompt engineering is a systems discipline. Here's the toolkit I bring to every engagement.
Structured, versioned prompt systems with clear roles, constraints, and few-shot scaffolding built for maintainability.
Golden datasets, LLM-as-judge rubrics, and regression suites that catch quality drift before your users do.
Multi-step tool-using agents with guardrails, retries, and graceful degradation for production reliability.
Knowing when a better prompt beats a fine-tune — and building the retrieval layer when it doesn't.
Token-budget-aware prompt compression and model routing that cuts spend without cutting quality.
Adversarial testing for jailbreaks, prompt injection, and hallucination under production traffic patterns.
A few systems I've designed, tuned, and shipped to production.
Redesigned a support agent's prompt chain to cut escalations and hallucinated policy answers, using retrieval-grounded responses and strict output schemas.
Built a modular persona + style-guide prompt system that let non-technical editors generate on-brand copy across 9 product lines.
Designed a multi-agent pipeline for market research: planning, web retrieval, synthesis, and self-critique loops with human checkpoints.
Built an LLM-as-judge evaluation harness with clinician-reviewed rubrics to gate every prompt change before deployment.
Own prompt strategy and evaluation infrastructure for a suite of customer-facing AI products used by 2M+ end users.
Shipped GPT-3-powered writing tools; built the company's first internal prompt-versioning and A/B testing system.
Built classical NLP pipelines and early transformer fine-tunes for document classification and entity extraction.
Thesis on discourse coherence modeling in generative dialogue systems.
"Aria doesn't just write prompts — she designs the whole reasoning process. Our hallucination rate dropped by half in one sprint."
Open to consulting engagements, fractional prompt-engineering roles, and interesting problems.
Email me →