In this post11 sections
- What candidates report about decomposition with data
- Read the table first: grain, keys and timestamps
- Find the user and the decision
- Propose one feature and one metric
- Write the first query out loud
- Worked example: fitness-app events
- Traps the columns hide
- Practice it before the real round
- Questions people ask
- Keep reading
- More from the blog
The interviewer shares a screen with a table on it, says “here’s some data, build something useful with it”, and goes quiet. You have a blank whiteboard, a column list you have never seen, and the urge to pitch an AI coach in the first minute. This post is the narrow version of our decomposition interview guide: what to do when the vague problem arrives as rows and columns.
The short answer: start from the columns, not from an idea. Read the table’s grain, keys and timestamps out loud. Then name one user and the decision the data can help them make, propose one feature and one metric, and write the first query that would prove the idea is worth building. The rest of this post works that method on a fitness-app events table, with the queries run and the traps marked.
What candidates report about decomposition with data
Two Palantir new-grad candidates on Blind said their decomposition interview came with data. Source 1Palantir Learning & Decomposition Interview (Blind)PublisherBlindSource typecandidate report on BlindSource 2Update: interview experience - Palantir new grad FDSE interview (Blind)PublisherBlindSource typecandidate report on Blind One of those candidates reported, in an update to an October 2022 post, being handed 8,000 records of London taxi data with 8 columns and asked for a first-cut solution for taxi drivers that could be built and deployed in a week. Source 1Palantir Learning & Decomposition Interview (Blind)PublisherBlindSource typecandidate report on BlindSource 2Update: interview experience - Palantir new grad FDSE interview (Blind)PublisherBlindSource typecandidate report on Blind That candidate, who was rejected, wrote on Blind that the asks were to suggest an idea, draw system components, come up with the APIs, and code if needed. Source 1Palantir Learning & Decomposition Interview (Blind)PublisherBlindSource typecandidate report on Blind
Palantir’s own careers page advises preparing by finding a public dataset and thinking of something interesting to do with it. Source 3Palantir Careers | Navigating Open-Ended QuestionsPublisherPalantir TechnologiesSource typecompany hiring page That is prep advice, not a description of the round. No employer publishes how often a decomposition prompt includes data, so prepare for both. For the taxi prompt itself, read our worked Palantir decomposition example.
The method below is ours, and it works at an AI product company too.
If the table is model-call logs
Grain: one row per request, or per retry? Keys: request_id against conversation_id, and is prompt_version reliable? Clocks: is latency measured on the client or the server? The open window: a thumbs-up can arrive minutes or days after the response, so recent requests look unrated. Hold them back until the rating window closes.
Read the table first: grain, keys and timestamps
Before you have an opinion, spend two or three minutes reading. Three questions, in this order.
The grain: what is one row? Say it in a sentence: “One row per event, per session, per user.” The grain decides every count you will make. If you think a row is a workout but it is an event, every number you give is inflated.
The keys: what joins to what? Name the column that identifies a row, and the columns that point at other things: a user, a session, a device, an order. Ask which keys are reliable. A key that repeats when it should not is your first data-quality finding.
The timestamps: when, and whose clock? Ask two things about every time column. Which time zone is it in? And is it when the thing happened, or when the server received it? A phone that syncs after a flight sends yesterday’s workout today.
One poster who had sat a Palantir interview wrote on Blind, in July 2022, that “clarifying the question and data schema/source is super important.” Source 2Update: interview experience - Palantir new grad FDSE interview (Blind)PublisherBlindSource typecandidate report on BlindSource 4Palantir FDSE Interview (Blind)PublisherBlind (Teamblind)Source typecandidate report on Blind Reading the columns out loud is how you do that without a script.
Words you can use:
“Before I pick a feature, I’d like to read the table. I think the grain is one row per event. Is
event_idunique, or can the client send the same event twice? And are these timestamps from the device clock or the server?”
Find the user and the decision
A table has no user in it. You have to find one. Look at the columns and ask: who would change what they do if they could see this data summarized well?
Name at least two candidates, one sentence each, then pick one. The pick should be a person with a decision that recurs, not a vague goal. “The content lead deciding which workout to rework next quarter” beats “improve engagement”, because a decision gives you a metric and a stopping point.
In a Blind thread about a Palantir data-focused open-ended round, a commenter who did not name their role wrote, in January 2025, that the round starts from a random feature and moves through the P0 features, technical decisions, data models and databases, what to analyze, which KPIs to track and what to present to executives. Source 5Palantir FDSE InterviewPublisherBlind (teamblind.com)Source typecandidate report on BlindSource 6Palantir FDSE Decomp/System Design InterviewPublisherBlind (teamblind.com)Source typecandidate report on Blind Picking a user and a decision early is what keeps that whole chain pointed at one thing.
Propose one feature and one metric
Now say the feature in one sentence, and say what it is not. Then define the metric precisely enough that two people would compute the same number:
- The event that counts. Which
event_typemeans success? - The denominator. Out of what? Sessions started, members active, orders placed?
- The window. Within the session, within
48h, within the calendar week? - What you exclude. Test accounts, sessions whose window has not closed, rows you cannot trust.
Palantir’s page on open-ended questions says to articulate the alternatives and trade-offs, arrive at a concrete approach, and deliver a functioning idea first, then expand it. Source 3Palantir Careers | Navigating Open-Ended QuestionsPublisherPalantir TechnologiesSource typecompany hiring page One feature and one metric is the functioning idea. For more on choosing the number and its guardrail, read our post on success metrics in case interviews.
Write the first query out loud
The first query is not the feature. It is the cheapest test of whether the feature is worth building. If the metric barely varies across the thing you want to act on, there is no decision to make, and you should say so and pick another feature.
Ask before you type: “Would you like me to write this in SQL, or sketch it in pseudocode?” Then narrate as you go: what the CTE builds, what one row of it is, what you are excluding and why. If window functions come up, our post on SQL window functions for FDE interviews covers the patterns and how to explain them.
Say the grain of every intermediate table
Each time you write a CTE, say what one row of it is: “one row per session”. It catches fan-out bugs before they happen, and it shows the interviewer you are checking your own work.
Worked example: fitness-app events
Here is a fictional table from a fitness app. The prompt: “This is a week of our event data. Build something useful.”
| Column | What it holds |
|---|---|
event_id | ID the app gives each event |
session_id | One workout; empty for app_open |
user_id | The member |
event_type | app_open, workout_started and so on |
workout_type | run, strength, cycling, yoga |
platform | ios, android or watch |
occurred_at | Device clock time |
received_at | Server time, in UTC |
duration_s | Seconds, on the finishing event |
The full fictional sample has 39 rows; here are six:
event event_type platform
e02 workout_started ios
e03 workout_completed ios
e03 workout_completed ios
e06 workout_started android
e13 workout_started watch
e14 workout_saved watch
Step 1: the grain and the keys. “One row per event. But e03 appears twice: same occurred_at (07:34:02), with received_at 07:34:03 and 07:36:40. So the client probably retried the send. I’ll ask how it retries, and check how often:”
SELECT
COUNT(*) AS n_rows,
COUNT(DISTINCT event_id)
AS n_events
FROM events;
On the sample it returns 39 rows and 38 distinct events. So “one row per event” is false, and anything that counts rows must dedupe first. session_id is the key that joins a start to its finish, which is why this table can support a completion metric at all.
Step 2: the event vocabulary, by platform. “Before I trust event_type, I want to see every value it takes on every platform.”
SELECT platform, event_type,
COUNT(*) AS n
FROM events
GROUP BY platform, event_type
ORDER BY platform, event_type;
The watch never sends workout_completed. It sends workout_saved. Say it and ask: “Does workout_saved on the watch mean the same as workout_completed on the phone?” Assume the interviewer says yes.
Step 3: the timestamps. “received_at is UTC. Does every occurred_at carry a zone, either a Z or an offset like +01:00?”
SELECT platform,
COUNT(*) AS no_zone
FROM events
WHERE occurred_at NOT LIKE '%Z'
GROUP BY platform;
Two Android rows have no zone, and their occurred_at runs an hour ahead of received_at: device-local time. On a real table I’d also allow offsets; this sample has none. Anything by hour of day or by date has to convert first, or those rows land in the wrong bucket.
Step 4: the user and the decision. “I see three users for this table. A member deciding whether to work out today, which points to a reminder. A support engineer asking which platform loses events. And the content lead deciding which workout type to rework first. I’ll take the content lead: the decision recurs every quarter, and session_id lets me measure it directly.” The member reminder is the feature worked in our fitness-app feature question, so practice that one next.
Step 5: one feature, one metric. “The feature is a weekly finish-rate table by workout type, with a drill-down by platform, that the content lead reads before planning. It is not a recommendation engine. The metric is finish rate: of sessions with a workout_started, the share with a workout_completed or workout_saved.”
Step 6: the first query.
WITH s AS (
SELECT session_id,
MAX(workout_type)
AS workout_type,
MAX(user_id) AS user_id,
MAX(CASE WHEN event_type =
'workout_started'
THEN 1 ELSE 0 END)
AS started,
MAX(CASE WHEN event_type IN (
'workout_completed',
'workout_saved')
THEN 1 ELSE 0 END)
AS finished
FROM events
WHERE session_id IS NOT NULL
GROUP BY session_id
)
SELECT workout_type,
COUNT(*) AS sessions,
COUNT(DISTINCT user_id)
AS members,
SUM(finished) AS finished,
ROUND(1.0 * SUM(finished)
/ COUNT(*), 2) AS rate
FROM s
WHERE started = 1
GROUP BY workout_type
ORDER BY rate;
Narrate it: “The CTE is one row per session, so the duplicate e03 counts once. I keep only sessions with a start, because the metric’s denominator is sessions started. s21 has a finish and no start: that’s a lost event I’d flag, not a workout I’d count. Without this filter, cycling reads 0.8.” On the sample, strength finishes at 0.5, cycling at 0.75, run at 0.83 and yoga at 1.0.
“A last check before I trust these: s16 started at 17:30 on the last evening, one second before the file ends. Its window hasn’t closed, so I drop sessions started within 2h of the cutoff. Cycling goes to 1.0. Strength is still lowest.”
Then read the members column: “Strength has six sessions from three members. Three misses, two members, and two of them are one Android member. That’s too thin to rework a workout type on, so the page shows members beside every rate.”
Step 7: the first version. “Week one: this query, with the cutoff, as a nightly job into a finish_rate_weekly table, and one page the content lead opens on Monday. Week two: the platform drill-down and the time zone fix.” That is a : thin, real, end to end.
Traps the columns hide
These are the ones that sink a dataset round, and the fix for each.
Tripwire: Counting rows as if they were things
You count events and call them workouts. Fix: say the grain, then count distinct keys. On the sample, counting rows would have counted the retried e03 twice.
Tripwire: Trusting an event name across platforms
Without the workout_saved check, strength finishes at 0.17 instead of 0.5, and yoga drops too. You would have sent the content lead after a logging gap. Fix: list every event value by platform or app version before any metric.
Tripwire: Mixing device time and server time
Local timestamps with no zone put evening workouts on the wrong day. Late syncs put a Tuesday workout in Thursday’s data. Fix: decide which clock each metric uses, convert to UTC, and filter on occurred_at, not received_at, when you mean “when it happened”.
Tripwire: Measuring sessions that have not finished yet
A session started five minutes before the data was pulled looks abandoned. On the sample, keeping s16 made cycling look like the second-worst workout. Fix: drop sessions whose window has not closed, and say the cutoff out loud.
Tripwire: Letting one member be the whole signal
In the sample, two of the unfinished strength sessions belong to one member. Fix: show a count of distinct members beside every rate, and say how much data the rate rests on.
A Blind user who said they used to work at Palantir listed, in March 2025, what a bad decomposition looks like: not asking questions, making large assumptions without clarifying, misunderstanding the problem before jumping in, impractical solutions, and not expanding on the v0. Source 2Update: interview experience - Palantir new grad FDSE interview (Blind)PublisherBlindSource typecandidate report on BlindSource 4Palantir FDSE Interview (Blind)PublisherBlind (Teamblind)Source typecandidate report on Blind Every trap above is a large assumption made about a column. Reading the table first is how you ask the questions instead.
Your first five minutes with a dataset
- Say the grain in one sentence, then check it with a distinct count
- Name the keys and ask which ones repeat or go missing
- Ask whose clock each timestamp uses, and whether any lack a zone
- List every value of the event or status column, by source
- Name two or three users and pick one with a recurring decision
- Say one feature, what it is not, and one metric with its window
- Ask whether they want SQL or pseudocode, then write the first query
Practice it before the real round
Reading this is not the same as doing it with someone waiting. Three ways to practice:
- Work the fitness-app feature question with a timer, then compare your answer with the model answer.
- Drill the queries that come up in dataset rounds: grouping page views into sessions and return rate by cohort.
- Get Pro for our lesson on dataset-first prompts, which covers entities and keys, the first endpoint and a live function inside a short screen. Pro starts with a 7-day free trial. The details are on the pricing page.
Then practice the asking half live in the free AI practice case. The free case gives you a city’s permit status history and an intake spreadsheet, and the AI customer reveals what is in them only when you ask. Ask for the grain, the key and how fresh each one is before you propose anything.
Questions people ask
Have candidates reported FDE decomposition interviews with a dataset?
Yes. Two Palantir new-grad candidates reported on Blind that their decomposition round came with data; one of them described London taxi records and a request for a first-cut solution for drivers. No employer publishes how often a dataset is included.Source 1Palantir Learning & Decomposition Interview (Blind)PublisherBlindSource typecandidate report on BlindSource 2Update: interview experience - Palantir new grad FDSE interview (Blind)PublisherBlindSource typecandidate report on Blind
What should I look at first when they hand me a table?
The grain (what one row is), the keys that join it to anything else, and the timestamps: which time zone, and whether a time is when the event happened or when it arrived. Those decide which features and metrics the table can support.
Has a candidate reported a decomposition round with data that asked for code?
Yes. One Palantir new-grad candidate reported on Blind, in an update to an October 2022 post, that the prompt asked them to suggest an idea, draw system components, come up with the APIs and code if needed; they were rejected. Ask the interviewer before you start writing.Source 1Palantir Learning & Decomposition Interview (Blind)PublisherBlindSource typecandidate report on Blind
Keep reading
Lessons
More from the blog
Interview rounds
‘We want an AI agent’: decomposing a vague AI request into a first version
A customer asks for AI agents everywhere. Cut it to one workflow, an eval set, human review and a launch criterion, worked on an insurer’s claims request.
Interview rounds
Choosing a success metric in a case interview: the number, its guardrail and who owns it
‘How would you measure success?’ exposes vague answers. Name the outcome metric, the guardrail and the person who will read it, with three worked cases.
Interview rounds
Palantir decomposition interview example: a full walkthrough of a reported prompt
A full walkthrough of a Palantir decomposition prompt one candidate reported, built on London taxi data, with what to say out loud at each step.