AI Engineer
Ship LLM features, not just prompts.
The complete applied-AI path: what the role actually is, the model landscape, working with model APIs and tokens, prompt engineering, safety, open-source models, embeddings and vector databases, RAG, agents, multimodal AI and the modern AI dev toolchain.
Saved on this device - no account needed.
- 1
Introduction
Know the role, the vocabulary and the landscape.
1 week0/5An AI engineer builds products on top of foundation models - integrating, grounding and evaluating them - rather than training models from scratch. It's a software engineering role with a new primitive, reachable from frontend, backend or full-stack.
Free resources
ML engineers train and deploy models from data; AI engineers compose pre-trained models into products. Different toolchains, different failure modes - and different interviews.
Prompt and context design, API integration, retrieval pipelines, evals, cost/latency budgets and safety review - the day-to-day surface of the job, and how it changes product development itself.
LLM, inference vs training, tokens, context window, embeddings, vector database, RAG, agent, AGI vs today's AI - the working glossary you'll use in every design discussion.
Free resources
Next-token prediction, sampling and temperature - and why hallucination is a property of the mechanism, not a bug to be patched.
Free resources
- 2
Pre-trained Models
Choose models like an engineer, not a fan.
1–2 weeks0/4Training frontier models costs millions; using them costs cents. The benefits (capability now, no ML team) and the trade-offs (cost per call, data governance, vendor dependency).
Anthropic's Claude, OpenAI's GPT models, Google's Gemini, Mistral, Cohere and Llama-family open models - plus cloud platforms (Azure AI, AWS Bedrock/SageMaker) that host them.
Free resources
Context length, knowledge cut-off dates, modality support and speed tiers - the spec sheet that decides model fit before any benchmark does.
Hallucination, prompt sensitivity, knowledge staleness and non-determinism - designing products that stay useful when the model is wrong.
- 3
Working with Model APIs
Call models reliably, safely, affordably.
2–3 weeks0/6System, user and assistant roles, multi-turn state and stop conditions - the request anatomy shared by every provider's API.
Free resources
Provider workbenches for iterating on prompts before writing code - parameters, comparisons and shareable experiments.
How text becomes tokens, counting them programmatically, and max-token limits - the unit your budget and context window are denominated in.
Free resources
Input vs output pricing, prompt caching, batching and model routing (small model first) - AI features have a unit cost you must design, monitor and defend.
Server-sent tokens and progressive UI - the difference between an app that feels instant and one that feels broken.
Adjusting weights on your examples: strong for style and narrow tasks, usually beaten by better prompts + RAG for knowledge. Know the decision framework.
- 4
Prompt Engineering
Treat prompts as versioned artifacts, not vibes.
2 weeks0/4Role, task, constraints, examples, output format - a repeatable template that makes results predictable across inputs.
Free resources
Examples that teach the pattern and reasoning room that improves accuracy - and when each actually helps.
Free resources
Schema-constrained JSON plus validate-and-retry loops - downstream code should never parse prose.
Testing prompts against fixed input sets, diffing outputs and keeping prompts in git next to code.
Build: Structured extraction service
An endpoint that turns messy text (resumes, invoices) into validated JSON with retries and cost logging.
- 5
AI Safety & Ethics
Ship features you can defend to security review.
1–2 weeks0/5Untrusted content smuggling instructions into trusted context - the signature attack of LLM apps, and layered defences for anything that reads user or web content.
Models inherit their training data's biases - testing across demographics and use cases before your users do it for you.
PII in prompts, data retention policies, end-user IDs for abuse attribution and regional processing - the questions legal will ask.
Input/output filtering, moderation endpoints, constrained outputs and human-in-the-loop for consequential actions.
Free resources
Red-teaming your own app with jailbreak and injection suites before launch - finding the failure modes on your schedule, not Twitter's.
- 6
Open-Source AI
Run models you control.
2 weeks0/4Open-weights models trade some capability for control, privacy and unit economics - the decision matrix for regulated data and offline use.
Finding models by task, reading model cards and licenses, and the Hub as the app store of open ML.
Free resources
Inference endpoints, the transformers library and Transformers.js in the browser - from Hub page to running output.
Pulling and running quantised models locally with an OpenAI-compatible API - prototyping and private workloads on your own hardware.
Free resources
- 7
Embeddings & Vector Databases
Give models a searchable memory.
2–3 weeks0/5Text mapped to vectors where distance ≈ meaning - the primitive behind semantic search, recommendations, anomaly detection and classification.
Free resources
Semantic search over docs, 'similar items' recommendations, outlier detection and cheap zero-shot classification - one primitive, many products.
Hosted APIs vs open Sentence Transformers - dimension, speed and cost trade-offs, and why you must never mix models in one index.
Free resources
pgvector, Chroma, Pinecone, Qdrant, Weaviate, FAISS and friends - pick one, learn indexing and filtering deeply; the concepts transfer.
Free resources
Cosine similarity, top-k retrieval, metadata filters and hybrid (keyword + vector) search - precision work that decides RAG quality.
- 8
RAG
Ground models in data they weren't trained on.
2–3 weeks0/4Retrieval for knowledge (fresh, auditable, cheap to update); fine-tuning for behaviour - the trade-off every 'chat with our data' project must get right.
Chunk → embed → store → retrieve → generate with citations. Each step has failure modes (bad chunks, stale index, irrelevant top-k) and fixes.
Size, overlap and structure-aware splitting so retrieval returns answers, not noise.
Orchestration frameworks vs calling SDKs directly - what they abstract, what they hide, and when plain code is clearer.
Free resources
Build: Docs chatbot with citations
RAG over a real documentation set; every answer cites its sources and admits when nothing relevant exists.
- 9
AI Agents
Automate multi-step work with control you can defend.
2–3 weeks0/5Declaring functions the model can call, executing them and returning results - the primitive underneath every agent.
Free resources
Reason → act → observe cycles, stop conditions and budgets - and why an unbounded loop is an outage waiting to happen.
Deterministic pipelines beat autonomous agents for most tasks - choose the simplest structure that works, escalate autonomy only when needed.
Conversation windows, summarisation and external state - keeping long tasks coherent past the context limit.
The Model Context Protocol standardises connecting models to tools and data - the plumbing layer under modern agent apps.
Free resources
- 10
Multimodal AI
Beyond text: images, audio and speech.
1–2 weeks0/4Vision-capable models reading screenshots, documents and photos - OCR-free document extraction is a product category now.
Free resources
Text-to-image APIs and prompt technique - plus the licensing and safety questions that ship with generated media.
Transcription (Whisper-class) and voice synthesis - the interface layer for voice-first products.
Free resources
Combining vision, audio and text in one flow (e.g. meeting recording → transcript → summary → action items) with your framework of choice.
- 11
Evals & Production
Prove quality; run it live.
2 weeks0/4Golden sets, LLM-as-judge with spot-check calibration and regression evals in CI - you can't improve what you don't measure.
Tracing prompts, tokens, latency and tool calls in production - failures you can debug instead of shrug at.
Rate limits, timeouts and idempotent retries with backoff - network engineering applied to model calls.
AI code editors and agentic CLIs (Claude Code, Cursor, Copilot) - using AI to build AI, and knowing when to trust it.
Free resources
- 12
Ship & Prove
A deployed AI product and verified proof of skill.
2 weeks0/2Ship one AI feature end to end - RAG app, agent or copilot - deployed, evaluated and documented.
Take the OneRoadmap AI Engineer certification to verify the skills recruiters are now screening for.
Free resources