FDE Toolkit · Portfolio Blueprint

What project actually gets you hired as an FDE?

You are competing for a role you have not done yet, so the portfolio is standing in for experience you cannot have until someone hires you. That only works if the project survives the way a senior engineer reads it. Here is what they check, the four builds that cover the bar, and how to make each one look like your job rather than a tutorial.

About this guide

We put this together while designing the project work for our own Forward Deployed Engineer programme. It draws on the FDE job descriptions we track across the frontier labs and enterprise vendors, on how the interview loop actually runs, on the build assignments companies hand candidates as take-homes, and on what the career switches that worked had in common.

We then took it to people who do the hiring: practising FDEs, AI engineers, hiring managers and the instructors who teach this. What follows is what survived those conversations, written for someone weighing up whether forward-deployed engineering is a move they want to make.

Treat it as directional. Which projects are right for you depends on the roles you are targeting, the domain you already know and how much time you have, so the sensible next step is to work your actual plan through with an FDE career mentor or a technical mentor who has sat on the other side of these interviews.

Want IK to ship you a portfolio like this? Talk to a program adviser.Book a call →

How a reviewer actually reads your project

Before it is a portfolio it is a repo, and somebody senior is going to open it with about fifteen seconds of patience. Four things decide whether they keep reading.

01
Does it run from a clean clone?

They will clone it and follow your README, so it wants one command to set up and one to run. Commit a .env.example, and check that no credential ever made it into the history. If getting it running needs a message from you, that is a fail before anyone has read a line of your code.

Reads as weak Instructions that quietly assume your laptop.

02
Is there a real eval, or is it vibes?

The question that separates a shipped system from a demo. Build an eval set of 20–50 cases, and make it cover the easy path, the messy input, the out-of-scope question and the one your system should refuse outright. Match the metric to the task: citation accuracy, schema validity, tool-call correctness. Then commit the results, so a reviewer sees the number without having to run anything.

Reads as weak "It looks right."

03
What does it do when things go wrong?

Real inputs arrive malformed. The upstream API times out. The model answers confidently about something it has never seen. So a reviewer goes looking for bounded retries, for an explicit refusal path, and for logs with enough in them to reconstruct what happened: request id, model, tool calls, latency, tokens, cost. Local JSONL is fine as long as a human can read it.

Reads as weak A happy path and a try/except that swallows the error.

04
Can you defend the decisions?

Every choice in the repo is a question waiting to be asked. Why that chunk size, why that model, why you did not fine-tune. What they want back is the trade-off you made and what it cost you, which is a very different answer from a tour of your dependency list.

Reads as weak Naming LangChain and Pinecone instead of what you decided and why.

What your GitHub repo should contain

The files a reviewer expects to find when they open the folder, and what belongs in each one. Most of these take an hour and are the difference between a repo that gets read and one that gets closed.

README.md

One paragraph on what it does. A demo, whether that is a screenshot, a GIF or sample output. Four commands: setup, run, test, eval. Then an honest note on what it does badly. The first screen has to be useful on its own.

src/

Modules named for what they do. Not one 900-line main.py, and not a notebook pretending to be a service.

tests/

Runnable with one command. Cover the deterministic parts a model cannot rescue: parsing, validation, routing, schema enforcement.

evals/

The dataset and the harness. results/ holds timestamped runs, committed, so the score is in the repo rather than in your memory.

ARCHITECTURE.md

A diagram, the constraint you were designing against, and the approaches you rejected. This is the file that survives the deep-dive.

.env.example

Every variable named, no values. The real .env never goes in.

Makefile / compose.yaml

So the whole setup is make run or docker compose up. Every minute a reviewer spends on setup is a minute not spent on your code.

The questions an interviewer asks about your project

Once a project is on your CV, one of the rounds is spent taking it apart. The hiring manager, and often a senior engineer sitting in, will work through most of the list below on whichever project interests them. Prepare answers for your strongest one.

  1. Walk me through this end to end.
  2. What business problem was it, and who was the customer?
  3. What was actually yours to decide here?
  4. Why this approach over the alternatives? Are you still comfortable with that trade-off?
  5. Is there an actual eval framework here, or is it vibes-based?
  6. How do you know it worked? What did you track?
  7. What was harder than you expected?
  8. How does this change at ten times the scale, or with dirtier data?
  9. What would you do differently starting again?

How to show work you are not allowed to make public

This is the most common worry senior engineers raise with us, and a fair one. Your best work is usually your paid work, and it sits in a private company repo you cannot link to. You do not need to. What an interviewer wants from a project is the thinking behind it, and there are four ways to hand them that without shipping a single line of your employer's code.

Publish the architecture, not the code

A design doc with the diagram, the constraint you were under, and the options you rejected. It carries no proprietary data and it is the artefact the deep-dive actually probes.

Rebuild the hard part on public data

Same problem shape, open or synthetic corpus. You keep the engineering and leave the employer's data where it is.

Publish the numbers, not the inputs

Latency, cost per call, eval deltas, adoption. Clear it with your employer first, then the result stands on its own.

Bring it to the room regardless

The probes above are about decisions, not source. You can answer every one of them about a repo nobody outside your company will ever see.

What to build: four projects, in this order

You do not need fifteen. Four cover every capability an FDE loop tests, and the fourth is the one an AI engineer would not already have. Each carries the take-home it pre-answers, so the work doubles as interview prep.

1

Grounded retrieval over a corpus that fights back

RAG & retrievalEvaluation

Answers that cite the document they came from, and a system that says it does not know when the corpus cannot answer. Most of the work lands in the retrieval and the access rules; the prompt is the small part.

The take-home it pre-answers

Build a retrieval assistant over this document set. Every answer must cite its source, and it has to refuse when the corpus cannot answer.

Make it yours
  • Healthcare: payer policy documents, where the wrong citation is a denied claim
  • Financial services: compliance manuals under an examiner who wants the source
  • Telecom: network runbooks that contradict each other across regions
  • Retail: supplier contracts with terms buried in amendments
  • Public sector: regulations where the version in force depends on the date
Build it with IK

P2 · SupportDesk-RAG

Build it yourself

Your team's runbooks or wiki. You already know which answers are wrong, which makes you the only person who can build the eval set.

2

An agent with real tools and a human gate

Agentic buildMulti-agent systems

Tool access against a database, a document store and a shell, with anything destructive stopped for human approval. The agent loop is the easy half. What you are really demonstrating is that the gate holds when the model goes looking for a way around it.

The take-home it pre-answers

Build an agent that can query a database, search documents and run shell commands. Anything destructive needs a human approval step. Show us how you tested that gate.

Make it yours
  • Ops / SRE: an incident assistant that reads logs and proposes, never executes, a rollback
  • Finance: AP matching where a payment release always needs a person
  • Support: ticket triage that can refund up to a cap and escalates past it
  • Sales ops: CRM hygiene where a merge or delete is proposed, not applied
Build it with IK

P3 · Multi-Agent Travel Planner, then C4 · Support Assistant

Build it yourself

The manual workflow on your team that everyone complains about. Automate the reading, leave the deciding.

3

A deployment hardened for somebody else's environment

Production hardeningAI securityEnterprise integration

Tracing, evals in CI, PII redaction, cost ceilings, and a blast radius you bounded deliberately. This is where the AI-security round gets decided: untrusted retrieval, tool descriptions as an attack surface, and guardrails that live in the code. A prompt is only advice.

The take-home it pre-answers

Put a router in front of three model providers with caching and failover. It should hold at a hundred requests a second with p95 under two seconds. Show us your numbers.

Make it yours
  • Any regulated buyer: the deployment has to survive their security review, not just your laptop
  • Multi-tenant SaaS: per-tenant isolation across the index, the cache and any fine-tuned model
  • On-prem / air-gapped: the customer's data never leaves, so the model comes to the data
Build it with IK

P7 · Production-Ready Fintech Support Agent

Build it yourself

Take a build you already have and harden it. Same repo, a much better story.

4

One full customer engagement, start to handover The one that makes it FDE

Full customer engagementSolution architecture

Discovery, scope, build, deploy, handover, with a real person on the other side who wanted something and can say whether they got it. The first three prove you can build. This one proves you can be put in front of a customer, which is the part an AI engineer's portfolio never has to answer for.

The take-home it pre-answers

Five hours and our public API. Build something a customer in this industry would actually find useful, then present it to us as though we were that customer.

Make it yours
  • Your own employer: the strongest version. A real stakeholder, a real constraint, and a result you can name
  • A team next to yours: someone whose workflow you understand but whose data you do not own yet
  • A small business or non-profit: when your employer is not an option, a real user with a real problem still counts
Build it with IK

C7 · PriorAuth AI, a five-week engagement at a hospital network

Build it yourself

Run it at your current employer, which is the version that counts as work rather than as a side project. The four steps for doing that are in the section above.

Every one of these is a round on the interview page. See the FDE interview loop → · Built on a specific skill set. The skills roadmap →

Does a project you built yourself count as real experience?

Not in the way three years on the job counts, and any interviewer will tell you the same. But that is the wrong comparison to worry about. Your competition is everyone else making the same move, and most of them turn up with a certificate and nothing built.

What narrows the gap is a real user, a real constraint and something somebody depends on. All three are sitting at your current employer, which is where the strongest version of this gets built.

The part most people skip

Find someone who wants the thing to exist

A project with somebody waiting on it stops being something you did at the weekend. It has a stakeholder, a deadline and consequences, which is what makes it work rather than practice, and it is the only kind that answers the interview question about a customer-facing engagement. There are three places to find one, and only the first needs your employer to say yes.

Inside your team
The workflow everyone complains about

The lowest-friction option. You already know the edge cases, you know who is affected, and you do not need anyone's permission to start looking. The catch is that you are close enough to skip the discovery, which is the half FDE interviews probe hardest.

A twenty-percent project
Someone else's team, someone else's problem

Take a problem belonging to a team next to yours. You arrive as an outsider who has to ask what they actually do before writing anything, which is exactly the shape of the real job. Harder to get moving, and worth more when you talk about it.

Volunteer and non-profit
Real users, real constraints, no procurement

A well-worn route for people moving in from outside tech, and it is not a lesser version. A non-profit has messy data, a small budget, someone who needs the thing to work, and nobody to build it. You get a genuine stakeholder and a deployment you can talk about, without waiting for your employer to say yes.

However you get one, run it like this The fourth step is what separates this from a side project.
  1. 01
    Find the workflow people complain about

    Not the most interesting technical problem. The one your team loses hours to every week, where you already know the edge cases and who is affected.

  2. 02
    Get one person to sponsor it

    A manager or a team lead who wants the outcome and will say so later. This single step is what turns a side project into work, and it is the step people skip.

  3. 03
    Scope it like an engagement

    Write down what it will do, what it will not, what you need access to, and how you will both know it worked. A page is enough. It is also the artefact the interview asks about.

  4. 04
    Ship it, then hand it over

    Deploy it where the team actually works, watch what breaks, and write the handover so it survives without you. Handover is the part FDE interviews probe and almost nobody has done.

How Interview Kickstart builds this portfolio with you

One way to get the four built: let IK ship them to you. You do not start from a blank repo. The programme hands you eight guided builds covering the AI engineering the role runs on, then a capstone you own end to end, and a full customer engagement on top of it.

Live guided projects

P1–P8 · built with an instructor, live

You build these alongside an instructor in a live session, one component at a time, with the reasoning explained as it is written. The finished repo matters less than watching someone experienced make the calls, before you have to make them on your own.

P1

CRM Lead Qualifier Agent

Build your first LLM-powered agent from scratch: function calling as a reasoning engine, tools for domain lookup and lead scoring, and a working think–act–observe loop. It is the 0→1 every later build stands on, the moment a model stops answering and starts acting.

Built with
  • Function calling
  • tool use
  • the ReAct loop
Common across every build
  • Python
  • Git & GitHub
  • Docker
  • AWS

Capstone projects

pick one of C1–C6, then the FDE capstone C7

These you own. You get a brief and a mentor, and the code, the architecture and the trade-offs are yours. You pick one of the six agentic capstones, then everyone does C7, the FDE capstone, which runs as a full customer engagement rather than a build.

C1

Finnie · AI Finance Assistant

A six-agent system on LangGraph + RAG that delivers personalised investment guidance, portfolio analysis, goal planning, and tax education through one conversational interface. Six specialists coordinated into a single coherent financial co-pilot.

Built with
  • LangGraph
  • RAG
  • six-agent orchestration
Common across every build
  • Python
  • Git & GitHub
  • Docker
  • AWS
Why IK

You graduate with the portfolio already built and mentor-guided: eight guided builds plus an FDE capstone you own end to end.

The popular AI stack to build on

Hiring here is stack-agnostic. Practitioners who run FDE loops report hiring people whose background is Azure or LangGraph into shops that run none of it, because the agentic and LLM fundamentals transfer and no one platform is a requirement. So pick a stack and start. Low-friction to ship fast, engineering-heavy to prove the production judgment the FDE bar is set at.

Low-friction stackship this weekend
Engineering-heavy stackproduction-grade · the FDE bar
Model
Hosted frontier API: Claude or a GPT-class model. No infra, pay per token.
Frontier API for reasoning + a self-hosted open model (Qwen / Llama via vLLM); fine-tune with QLoRA when the domain needs it.
Language
Python.
Python + TypeScript.
AI coding assistant
Claude or Claude Code, pair-programming the whole build.
Claude Code in the loop for the tests and the eval harness, not just the code.
Frameworks
LangChain + LangGraph, or the OpenAI Agents SDK. Agents, RAG and orchestration out of the box.
LangGraph + custom MCP servers (FastMCP); OpenAI / Claude Agent SDK; Pydantic typed contracts.
Cloud & data
Managed vector store (Chroma or Pinecone serverless), nothing to run.
Azure AI Foundry / AWS Bedrock / Vertex AI; Docker + Kubernetes; Terraform; pgvector at scale.
Deployment
Vercel for the UI + serverless functions; Streamlit for a fast demo front-end.
Vercel front-end + containerized FastAPI; LangSmith tracing + DeepEval gates in CI; Guardrails / Presidio on the boundary.

Project ideas to start from

Two lists that are worth building from. The first is what companies actually hand candidates and then score. The second is what you can ship inside your own company, which is the shortest path to a project with a real stakeholder attached.

01
Cited RAG with a refusal path RAG

Retrieval over a supplied document set where every answer carries its source and unanswerable questions get a refusal, not a guess. Ship the eval set that proves both.

02
Fully open-source agentic RAG RAG

Same problem, no hosted API. Local model plus an open orchestration stack, which is what a customer who cannot send data outside their walls will ask for.

03
Tool-using agent with an approval gate Agent

Database, document and shell access, with destructive actions held for a human. What they are looking for is a test that tries to get through the gate.

04
Documents to typed JSON with confidence Pipeline

Scanned or messy documents into a strict schema, with a per-field confidence score and a stated line for what a human must check.

05
Multi-provider router Pipeline

Caching, failover and cost routing across three model providers against a stated throughput and p95 target. Report the numbers you actually hit.

06
Eval harness for an agent Eval

Quantitative scoring of an agent's output across correctness, refusal behaviour and reasoning quality, with the harness itself defensible under questioning.

07
Guardrail and PII layer Eval

Redaction across all four entry points (user input, retrieved context, tool results, model output) while keeping the answer readable.

08
Clone the repo, fix three issues Agent

Their codebase, their open issues, plus a recorded walkthrough of what you changed and why. Scored async on the code and on how you explain it.

09
Conversation to structured record Pipeline

Audio or transcript into clinical, meeting or CRM documentation with a schema that downstream systems can trust.

10
Custom MCP server for a real system MCP

Expose a live system's operations as tools with least-privilege scoping and an audit trail, so any model client can call it safely.

Why Interview Kickstart
25,000+
alumni network across tech
FAANG+
instructors from Google, AWS, Databricks, Microsoft and Meta
1:1
mentorship + FDE-tuned mock interviews
End-to-end
placement support
Build an FDE portfolio that gets interviews.

An Interview Kickstart advisor walks you through where you stand today, the exact gap to close, and the fastest route to a Forward Deployed Engineer offer, built around your background.

Book a call with an advisor →