An embedding is a fixed-length vector of numbers that a model produces for a piece of text (or an image, or a row), placed so that items with similar meaning end up close together under cosine similarity or dot product. Sentence-level text embeddings of the kind used for search were popularized by Reimers and Gurevych, Sentence-BERT, and the MTEB benchmark compares models across tasks.
In an FDE interview
Scale AI’s Frontier Agents postings, as of September 2026, ask for experience building or deploying AI-powered applications with LLM APIs, frameworks, , retrieval systems or vector databases. Source 1Frontier Agents Engineer (Forward Deployed Engineering)PublisherScale AI (Greenhouse)Source typecompany job posting If a coding exercise asks for top-k similarity over embeddings, normalize the document vectors once at index time and the query vector at query time, so scores = docs @ q is the cosine. Then take the top k with np.argpartition(-scores, k - 1)[:k] and sort only those (heapq.nlargest without NumPy), rather than sorting every score.
In a design discussion, name what embeddings are bad at: exact identifiers such as part numbers and error codes, and negation. Queries and documents must use the same model, with the query or document prefix (or input_type) the model expects. Some libraries cut text past the model’s input limit without an error, so chunk size follows the model. Changing the model means embedding the corpus into a second index, comparing the two on the same labeled queries, then switching reads. Choose a model by testing a shortlist on the customer’s own queries, not by leaderboard rank alone.
Top-k cosine without a vector database is the coding version.