A practice prompt we wrote. No company or candidate report names it, so it carries no company tag.

How to answer

This looks like a warm-up, but the answer turns on three definitions you have to make before you parse anything. Make them out loud.

  1. Ask for a real line. Get the format and the latency unit. If the interviewer leaves it to you, say your assumption: the common log format (nginx’s combined without the referer and user ) with $request_time appended by a custom log_format, in seconds.
  2. Define “endpoint”. Method plus path, query string removed, numeric, UUID and long hex segments collapsed to {id}. Say why: otherwise /orders/8812 and /orders/8813 are separate rows, each with one sample.
  3. Define the percentile. Name the method you use, for example nearest rank, and compute the rank in integer math. Say that an endpoint with very few requests has a p95 that is really its maximum, so you set a minimum sample count and list the thin endpoints separately.
  4. Parse defensively and count what you drop. A loose line regex, then a classifier with a counter of rejected lines by reason (bad request line, bad latency, unparsed) and a sample of each reason in the output. Never drop a line silently.
  5. Report, and ask what it’s for. Endpoint, count, p50, p95, max and total time, sorted by p95 with total time beside it, because a slow endpoint nobody calls matters less than a moderately slow one everyone calls. Offer a per-hour breakdown, because a whole-day p95 hides a lunchtime spike.
  6. Test with numbers you can check by hand. Twenty values from one to twenty: nearest-rank p95 is the nineteenth.

The trap is sorted_vals[int(0.95 * n)]. It lands one position too high whenever 0.95 * n is a whole number, so on twenty values it returns the maximum instead of the nineteenth. The second trap comes in the follow-up: averaging per-server p95 values. Percentiles don’t average; you merge the raw samples or mergeable histograms.

GlossaryAgentA system in which a model chooses steps and tool calls to complete a task, within limits the design sets.More on Agent

Follow-ups

What the interviewer may ask next, once your first answer is on the table.

  • Which percentile definition did you use, and does it matter here?
  • You have one log file per server. Can you combine each server’s p95?
  • The file is too large to hold every latency in memory. What changes?

Where answers go wrong

  • Treats every URL as its own endpoint, so each order ID becomes a row and the report ranks noise.

Answer this in two minutes

Write the answer you would say out loud. The clock starts with your first word.

Two minutes

Compare with the model answer

Model answer

“I’ll assume the server writes this format, which is nginx’s common log format plus $request_time. The nginx log module documentation defines $request_time as seconds with millisecond resolution, and the predefined combined format would add the referer and user after the byte count:”

log_format timed '$remote_addr - $remote_user [$time_local] "$request" '
                 '$status $body_bytes_sent $request_time';

203.0.113.9 - - [12/Mar/2026:10:15:32 +0000] "GET /api/orders/8812?expand=items HTTP/1.1" 200 1532 0.231

“An endpoint is the method and the path with IDs collapsed. I’ll use nearest-rank percentiles, because the result is always a latency that actually happened and it’s easy to check by hand. The line regex is deliberately loose, and a second step says why a line was rejected.”

import re
from collections import Counter, defaultdict

LINE = re.compile(
    r'^(?P<ip>\S+) \S+ \S+ \[(?P<ts>[^\]]+)\] '
    r'"(?P<request>[^"]*)" '
    r'(?P<status>\d{3}) (?P<bytes>\d+|-) (?P<rt>\S+)$'
)
REQUEST = re.compile(
    r"^(?P<method>[A-Z]+) (?P<target>\S+) HTTP/[0-9.]+$"
)
LATENCY = re.compile(r"^\d+\.\d+$")
UUID = r"[0-9a-f]{8}(-[0-9a-f]{4}){3}-[0-9a-f]{12}"
ID = re.compile(rf"^(\d+|{UUID}|[0-9a-f]{{24,}})$", re.I)


def endpoint(method: str, target: str) -> str:
    path = target.split("?", 1)[0]
    parts = ("{id}" if ID.match(s) else s
             for s in path.split("/"))
    return method + " " + "/".join(parts)


def percentile(sorted_vals: list[float], p: int) -> float:
    """Nearest rank: the smallest value with at least
    p% of the samples at or below it."""
    # ceil(p * n / 100) in integer math, so no float error
    rank = (p * len(sorted_vals) + 99) // 100
    return sorted_vals[max(rank, 1) - 1]


def classify(line: str):
    """Return (endpoint, latency_ms) or (None, reason)."""
    m = LINE.match(line)
    if not m:
        return None, "unparsed"
    r = REQUEST.match(m["request"])
    if not r:
        # "-" (no request sent), or TLS bytes on a plain port
        return None, "bad_request_line"
    if not LATENCY.match(m["rt"]):
        return None, "bad_latency"
    ep = endpoint(r["method"], r["target"])
    return ep, float(m["rt"]) * 1000


def analyze(lines, min_count: int = 20, top: int = 10):
    latencies: dict[str, list[float]] = defaultdict(list)
    rejected: Counter[str] = Counter()
    examples: dict[str, str] = {}
    for line in lines:
        line = line.rstrip("\n")
        if not line.strip():
            rejected["blank"] += 1
            continue
        ep, value = classify(line)
        if ep is None:
            rejected[value] += 1
            examples.setdefault(value, line[:200])
            continue
        latencies[ep].append(value)

    rows = []
    for ep, ms in latencies.items():
        ms.sort()
        rows.append({
            "endpoint": ep, "count": len(ms),
            "p50_ms": percentile(ms, 50),
            "p95_ms": percentile(ms, 95),
            "max_ms": ms[-1], "total_s": sum(ms) / 1000,
        })
    enough = [r for r in rows if r["count"] >= min_count]
    ranked = sorted(enough, key=lambda r: r["p95_ms"],
                    reverse=True)[:top]
    too_few = [r["endpoint"] for r in rows
               if r["count"] < min_count]
    return ranked, too_few, rejected, examples


def report(ranked, too_few, rejected, examples) -> str:
    out = [f"{'endpoint':<24}{'count':>6}{'p50':>6}"
           f"{'p95':>6}{'max':>6}{'total_s':>8}"]
    for r in ranked:
        out.append(
            f"{r['endpoint']:<24}{r['count']:>6}"
            f"{r['p50_ms']:>6.0f}{r['p95_ms']:>6.0f}"
            f"{r['max_ms']:>6.0f}{r['total_s']:>8.1f}")
    out.append(f"too few samples: {', '.join(too_few)}")
    for reason, n in rejected.most_common():
        out.append(f"rejected {reason}: {n}")
        if reason in examples:
            out.append(f"  e.g. {examples[reason][:80]}")
    return "\n".join(out)

“Run over a day of sample logs, the report reads:”

endpoint                 count   p50   p95   max total_s
GET /admin/export           25  2766  5081  6736    77.9
GET /api/search            300   380  1438  3843   156.0
POST /api/checkout         900   267   556  1075   263.0
GET /api/orders/{id}      1200   109   252   641   148.3
too few samples: GET /api/debug
rejected bad_request_line: 3
  e.g. 198.51.100.4 - - [12/Mar/2026:10:15:32 +0000] "-" 400 0 0.000
rejected unparsed: 2
  e.g. 203.0.113.9 - - [12/Mar/2026:10:15:32 +0000] "GET /api/sea
rejected bad_latency: 1
  e.g. 203.0.113.9 - - [12/Mar/2026:10:15:32 +0000] "GET /api/search HTTP/1.1" 200 12 -

“The admin export tops the p95 list, but checkout cost users more total time, which is why both columns are there. analyze takes any iterable of lines, so I can pass an open file and it streams. The report prints the rejected counts and an example of each, because a format change that breaks the regex should show up as a large unparsed count, not as a quiet drop in traffic. The lines that actually break parsers are a request field of "-", logged when a client opens a connection and sends no request, and TLS handshake bytes sent to a plain-HTTP port; both land in bad_request_line.”

def test_nearest_rank():
    vals = [float(i) for i in range(1, 21)]
    assert percentile(vals, 95) == 19.0
    assert percentile(vals, 50) == 10.0
    assert percentile([5.0], 95) == 5.0


def test_ids_collapse_and_query_dropped():
    got = endpoint("GET", "/api/orders/8812?expand=items")
    assert got == "GET /api/orders/{id}"


def test_bad_lines_counted_not_dropped():
    ts = "[12/Mar/2026:10:15:32 +0000]"
    good = f'1.2.3.4 - - {ts} "GET /a HTTP/1.1" 200 5 0.100'
    bad = f'1.2.3.4 - - {ts} "-" 400 0 0.000'
    ranked, _, rejected, examples = analyze([good] * 20 + [bad, ""])
    assert ranked[0]["count"] == 20
    assert rejected == {"bad_request_line": 1, "blank": 1}
    assert examples

“After the first real run I count distinct endpoints per leading path segment. If one prefix has far more rows than the API has routes, I missed an ID shape, such as a slug or a hash, and I add it before trusting the ranking.

Before I sort, I’d ask what the report will decide. If it’s where to spend engineering time, total_s matters as much as p95: an admin export that few people call can top the p95 list while checkout wastes far more of users’ time. I’d print both, and a per-hour p95 for the top rows so a spike isn’t averaged into a whole day.

Then I’d check what the latency actually measures. The nginx log module documentation says $request_time runs from the first bytes read from the client until the log is written after the last bytes are sent, so a slow mobile connection or a long streamed response inflates it. $upstream_response_time, from the upstream module, is the time spent receiving the response from the backend, so logging both separates a slow backend from a slow client. For an LLM endpoint that streams tokens, I’d ask whether ‘slow’ means time to first token or total time, because a whole-response p95 mostly measures how long the answers are.

Which requests count is a choice I’d state rather than make silently. I keep every status code, because a request that failed slowly was still slow for the user. I’d add a flag to split out server errors (5xx), and I’d look hard before excluding status 499, which nginx logs when the client closed the connection before the response: those are often requests the client gave up waiting for, and dropping them makes a slow endpoint look healthy.

On definitions: nearest rank and linear interpolation, which is NumPy’s default, differ a little on small samples and converge on large ones. What matters more is that the report says which one it used.

On multiple servers: I can’t average their p95 values, because a percentile of the whole isn’t a mean of the parts’ percentiles. I’d concatenate the samples, or keep a histogram per endpoint per server and merge those.

On memory: I store one float per request, which is fine for a day of logs on one box. Past that I’d keep a bounded summary per endpoint: an HdrHistogram, whose precision you set in significant digits, or a t-digest. Memory stays bounded. HdrHistogram’s relative error is fixed by that precision, while t-digest is accurate at the tails in practice but has no guaranteed bound. Both merge across servers, which answers the previous question too.

If the logs already sit in a warehouse, with the IDs collapsed into an endpoint column on load, it’s one query. This is Databricks or Spark SQL; Snowflake calls the function APPROX_PERCENTILE:”

SELECT endpoint,
       count(*) AS n,
       percentile_approx(latency_ms, 0.95) AS p95_ms
FROM requests
GROUP BY endpoint
HAVING count(*) >= 20
ORDER BY p95_ms DESC;

“Both functions are approximate, so the report says that, for the same reason it names the percentile method.”

Next, in Pro

In Pro, the trip summaries question takes the streaming idea further: events arrive out of order, state has to be bounded, and each result is emitted when a trip ends.

GlossaryAgentA system in which a model chooses steps and tool calls to complete a task, within limits the design sets.More on Agent