MLOps for Large Language Models - RAG pipelines, vector databases, prompt management, and cost optimization.
LLMOps applies MLOps principles to Large Language Models. LLMs have unique challenges: massive size, high cost, hard to evaluate, and can hallucinate.
Traditional ML: Train a custom model from scratch
LLMOps: Use a pre-trained model, customize via prompts, RAG, or fine-tuning
# rag_pipeline.py
from langchain_community.vectorstores import Chroma
from langchain_community.embeddings import OpenAIEmbeddings
from langchain.text_splitter import RecursiveCharacterTextSplitter
from langchain.chains import RetrievalQA
from langchain_anthropic import ChatAnthropic
class RAGPipeline:
def __init__(self, docs_dir):
self.embeddings = OpenAIEmbeddings()
self.llm = ChatAnthropic(model="claude-sonnet-4-20250514")
def ingest(self, documents):
splitter = RecursiveCharacterTextSplitter(
chunk_size=1000, chunk_overlap=200)
chunks = splitter.split_documents(documents)
self.vectorstore = Chroma.from_documents(
chunks, self.embeddings,
persist_directory="./chroma_db")
print(f"Ingested {len(chunks)} chunks")
def query(self, question):
retriever = self.vectorstore.as_retriever(
search_type="mmr", search_kwargs={"k": 4})
qa = RetrievalQA.from_chain_type(
llm=self.llm, retriever=retriever,
return_source_documents=True)
return qa.invoke({"query": question})| Database | Type | Best For |
|---|---|---|
| pgvector | PostgreSQL extension | You know PostgreSQL! |
| Pinecone | Managed SaaS | Production, no-ops |
| Weaviate | Open-source | Hybrid search |
| ChromaDB | Open-source | Prototyping |
| Qdrant | Open-source | High performance |
| Strategy | Savings | How |
|---|---|---|
| Response caching | 50-80% | Cache similar queries with Redis |
| Model routing | 40-60% | Simple queries to cheap models, complex to expensive |
| Prompt optimization | 20-40% | Shorter prompts = fewer tokens |
| Batch processing | 30-50% | Use batch APIs for non-urgent tasks |
RAG systems are fundamentally data pipelines: ingest > transform/chunk > embed > store > retrieve > serve. This is ETL with vectors instead of tables.