Vectorization Expert
Overview
As a specialist in scalable retrieval systems, you will design and implement end-to-end vectorization pipelines that turn text, PDFs, and multimodal data into high-quality embeddings. You’ll optimize multi-database vector stores and advanced indexing to enable fast, accurate search across large corpora. You will shape hybrid search workflows and robust evaluation frameworks to measure performance and faithfulness. You’ll collaborate with engineering and product teams to deliver high-perf retrieval solutions that scale with the business.
ResponsibilitiesLead design and development of scalable vectorization pipelines to convert text, PDFs, and multimodal data into embeddings using models like BGE, Ada, or CohereManage and optimize vector databases (Pinecone, Weaviate, Milvus, pgvector), including building efficient indexing strategies (HNSW, DiskANN)Develop advanced chunking and parsing techniques to maintain context and improve retrieval accuracyBuild and fine-tune hybrid search workflows combining dense vector search with sparse keyword methods (BM25)Create evaluation frameworks and gold-standard datasets to measure retrieval performance (Hit Rate, MRR, Faithfulness)Implement re-ranking pipelines using cross-encoders to refine results before sending them to LLMsWork closely with engineering and product teams to ensure scalable, high-performance retrieval systems
Key requirementsteam collaborationcross-functional communicationproblem-solvingvectorization pipelinesembeddingsBGE