Where it comes from
- Reported at PalantirSource 1Update: interview experience - Palantir new grad FDSE interview (Blind)PublisherBlind (Teamblind)Source typecandidate report on BlindSource 2Palantir Decomposition Interview Tips (Blind)PublisherBlind (Teamblind)Source typecandidate report on BlindSource 3Palantir Deployment Strategist Interview (Blind)PublisherBlind (Teamblind)Source typecandidate report on Blind More Palantir questions The Palantir interview guide
How to answer
This is the table the interviewer hands you: a few events from a few users, each an app open, a workout start or a workout completion, with a timestamp, a device, a workout type and a duration. It has six columns, so on a phone, swipe it sideways. Read all of it before you scroll on: some rows are wrong on purpose.
user_id|event_type |occurred_at |device_id |workout_type|duration_s
4117 |app_open |2026-03-02T06:58:10Z|ios-a91 | |
4117 |workout_started |2026-03-02T07:01:44Z|ios-a91 |run |
4117 |workout_completed|2026-03-02T07:34:02Z|ios-a91 |run |1938
4117 |workout_completed|2026-03-02T07:34:02Z|ios-a91 |run |1938
5230 |workout_started |2026-03-02T18:20:05Z|watch-7c |strength |
5230 |app_open |2026-03-03T12:02:47Z|watch-7c | |
6802 |workout_completed|2026-03-03 19:45:00 |android-3f|cycling |2710
6802 |app_open |2026-03-04T08:10:31Z|android-3f| |
One candidate reported on Blind, in August 2022, that in the of the onsite for a Palantir new grad role, “They give a sample data set and ask how can you use this data to do something.” Source 1Update: interview experience - Palantir new grad FDSE interview (Blind)PublisherBlind (Teamblind)Source typecandidate report on BlindSource 2Palantir Decomposition Interview Tips (Blind)PublisherBlind (Teamblind)Source typecandidate report on BlindSource 3Palantir Deployment Strategist Interview (Blind)PublisherBlind (Teamblind)Source typecandidate report on Blind This question is that shape. The dataset is the prompt, so the answer has to end in a table, an endpoint and a number you can compute. Open at the level of the data model and the metric, and write the query when the interviewer asks for it. Our method runs the moves in this order:
- Read the table before you pick anything. Say the grain out loud: “one row per event, per user, per device”. Name the columns you trust and the ones you doubt, such as a timestamp with no timezone or an event that fired twice.
- Name two or three candidate features in a sentence each, then let the data pick. Say the check that would decide: “Of members who miss their usual day, how many come back within the week on their own? If nearly all do, a reminder adds nothing.” Palantir’s careers page on open-ended questions says to articulate the alternatives and trade-offs, then arrive at a concrete approach and deliver a functioning idea first. Source 4Palantir Careers | Navigating Open-Ended QuestionsPublisherPalantir TechnologiesSource typecompany hiring page
- Pick one user and one decision. “The member deciding whether to work out today” beats “engagement”. Give the feature in one sentence and say what it is not.
- Draw the data model. Name the entities and keys, and the derived table your feature reads. The raw events are the input, never the thing you serve.
- Put the work where it belongs, then name the first endpoint. A batch job does the scan. The live path reads a precomputed row or makes one indexed lookup, and every endpoint has a real caller.
- Define the metric and how you will know it moved. Name the event that counts as success, the window and a holdout group.
- Cut to two weeks. The first week is a thin version on real data that one real user sees. The second week adds the metric, the holdout and the dirty-data fixes. Say what you fake.
The trap is a clever feature, such as an AI coach or a social feed, that eats the whole round and never reaches a key, a query or a metric. Stay plain and finish.
Follow-ups
What the interviewer may ask next, once your first answer is on the table.
- Which columns does your feature need, and which are missing or dirty?
- The table grows by hundreds of millions of rows a day. What does your first endpoint return?
- Write the query or function that computes your metric.
Where answers go wrong
- Picks a clever feature and never gets to a data model or a metric.
- Serves the feature from a scan of raw events instead of a precomputed row.
- Measures before and after instead of against a holdout.
Answer this in two minutes
Write the answer you would say out loud. The clock starts with your first word.
Compare with the model answer
Model answer
Reading the table. “Before I pick, let me read it. The grain is one row per event, per user, per device. The two completed runs for 4117 are the same workout with the same timestamp, most likely a phone that retried a sync after being offline, so I’ll ask how the client retries and dedupe on (user_id, device_id, event_type, occurred_at). Every timestamp ends in Z except the one from android-3f, so at least one client writes device-local time. I’ll ask which, and until I know I read a view, events_clean, that dedupes and converts occurred_at to UTC. And 5230 started a strength session that never completed: an abandoned workout or a lost event.”
Options, then a pick. “Three options: a comeback reminder, a weekly summary, or flagging workouts that start and never complete. I’d run one query first: of members who miss their usual weekday, how many complete a workout in the next seven days without help? If that number is low, a reminder has room to work. If it’s high, I’d pick the drop-off flag instead. Assume it’s low. The user is a member who works out on a regular weekday and misses it. The decision is whether to send one push today, and when. It is not a coach, a plan generator or a streak badge.”
Data model. “One derived table, rebuilt nightly from events_clean:”
CREATE TABLE member_habit (
user_id BIGINT PRIMARY KEY,
usual_dow SMALLINT, -- weekday with the most completed workouts, last 8 weeks
usual_hour SMALLINT, -- median local start hour on that weekday
last_completed TIMESTAMPTZ,
tz TEXT -- from the user profile; NULL means skip the member
);
“Members with a null timezone or fewer than 4 workouts in the window get no reminder. That is a deliberate cut, not a bug.”
Where the work happens. “The nightly job scans only the day’s partition of events, adds it to a small table of completed workouts per member, weekday and local start hour, rebuilds member_habit from the last eight weeks of that table, and records today’s decision for every member it considered, in both arms, in reminder_decisions(user_id, decided_at, arm, would_send, sent, reason). Only the treatment arm’s sends go on reminder_queue(user_id, send_at, reason). Logging the holdout’s decisions, the reminders we would have sent, is what makes the comparison fair: it gives me the same kind of member on both sides. At send_at, the push worker checks one thing against live data: an indexed lookup on events(user_id, occurred_at) for a workout_completed since the member’s local midnight. If there is one, it drops the reminder, so nobody who ran that morning gets nagged that evening; otherwise it sends and sets sent on the decision row. That lookup is the only raw-event read on the hot path, and it hits an index, not a scan. The first endpoint is GET /members/{id}/reminders for support and debugging: the recent decisions for one member, and why each was made.”
Metric. “Success is a workout_completed within 48h of the decision. I hold out a slice of eligible members who get no reminder, because the comparison that matters is reminded against not reminded, not before against after:”
SELECT r.arm,
COUNT(*) AS eligible,
AVG((EXISTS (
SELECT 1 FROM events_clean e
WHERE e.user_id = r.user_id
AND e.event_type = 'workout_completed'
AND e.occurred_at >= r.decided_at
AND e.occurred_at < r.decided_at + INTERVAL '48 hours'
))::int) AS return_rate
FROM reminder_decisions r
WHERE r.would_send
AND r.decided_at < now() - INTERVAL '48 hours'
GROUP BY r.arm;
“The arm is fixed per member, by a stable hash, for example hashtext(user_id::text) % 10 = 0 in Postgres as the holdout, so nobody switches sides. would_send is true in both arms when the member qualified, so members who were never candidates stay out. I drop the last two days because their window hasn’t closed. I’d also watch the opt-out rate, because a reminder that wins workouts but loses push permission is a loss. A holdout of one member in ten is small, so before launch I write down the smallest lift worth shipping and run until the holdout is big enough to see it, and I don’t stop on the first good day.”
Two-week cut. “In week one I build events_clean, the nightly job, the queue and the worker’s check, and send reminders to internal staff accounts through your existing notification service. In week two I open it to the holdout test, fix the timezone and dedupe rules on what staff saw, and ship the metric query as a dashboard. What I am deferring: a learned send time, multiple reminders and anything that reads workout content.”