Skip to main content

Command Palette

Search for a command to run...

Understanding Vector Search

Updated
•3 min read•View as Markdown
Understanding Vector Search
J
2nd Year IT Student | AI Agents • Open Source • DevOps • MLOps • Cloud Infrastructure | Building scalable, reliable systems.

How Data Becomes Searchable: The Embedding Pipeline

Before vector search can work, raw data needs to be converted into a numerical format that machines can compare.

This process is called embedding generation.

The embedding pipeline looks like this:

              Raw Data
                 │
                 ▼
        ┌─────────────────┐
        │ Embedding Model │
        └─────────────────┘
                 │
                 ▼
        ┌─────────────────┐
        │ Embedding Vector│
        │ [0.3, 0.1, ...] │
        └─────────────────┘
                 │
                 ▼
        Database with Vector
        Search Capability

Step 1: Input Data

The input data can be:

  • Text documents

  • Images

  • Audio

  • Videos

  • PDFs

  • Source code

Example:

"Introduction to Kubernetes"

Step 2: Generate Embeddings

The data is passed through a specialized machine learning model called an embedding model.

The model converts the information into a numerical representation:

"Introduction to Kubernetes"

            ↓

[0.23, -0.45, 0.87, ...]

These numbers capture the semantic characteristics of the data.


Step 3: Store Embeddings

The generated vector can be stored along with the original data.

For example, in MongoDB:

{
  "title": "Introduction to Kubernetes",
  "content": "A guide to containers and orchestration",
  "embedding": [0.23, -0.45, 0.87]
}

This allows existing data collections to become searchable by meaning.


Vector Search Workflow

When a user enters a query, the query goes through the same embedding process.

The system compares the query embedding with stored embeddings and retrieves the closest matches.

                 User Query
                     │
                     ▼
            ┌─────────────────┐
            │ Embedding Model │
            └─────────────────┘
                     │
                     ▼
              Query Vector
            [0.25, -0.41, 0.89]
                     │
                     ▼
        Compare with Stored Embeddings
                     │
                     ▼
          Similarity / Nearest Neighbor
                  Search
                     │
                     ▼
              Top-K Results

Instead of asking:

"Does this document contain the exact words from my query?"

vector search asks:

"Which stored information is closest in meaning to my query?"

This enables semantic search.


Consider a database containing:

Document 1:
How to train a puppy

Document 2:
Best food for dogs

Document 3:
Python programming basics

A user searches:

How do I raise a young dog?

The query is converted into an embedding.

The vector search system understands relationships such as:

young dog ≈ puppy

raise ≈ train

Therefore, it retrieves relevant results even without exact keyword matches.


Vector search usually returns the top K closest results.

For example:

Query:
"Best resources to learn AI"

k = 3

Results:

1. Machine Learning Roadmap
2. Deep Learning Fundamentals
3. AI Engineering Guide

Here, k = 3 means the system returns the three most similar results based on vector similarity.


MongoDB Vector Search Architecture

MongoDB stores embeddings alongside the original documents.

The workflow:

  1. Store documents in MongoDB.

  2. Generate embeddings using an embedding model.

  3. Add embeddings as fields inside existing documents.

  4. Create a vector search index.

  5. Convert user queries into query embeddings.

  6. Retrieve the most similar documents.

The important idea is that vector search extends existing databases instead of requiring completely separate storage systems.

MongoDB Document

{
  data: "...",
  metadata: "...",
  embedding: [...]
}

          │

          ▼

     Vector Index

          │

          ▼

  Similarity Search Results
11 views