Infinity Technologies
AI Startup
RemoteRemote OKFull-time$140k – $220k / year
About this role
Infinity Technologies
LLM / RAG Engineer
LLM / RAG Engineer
Middle
Senior
Role
LLM / RAG Engineer
Location
Remote / Hybrid
Technologies
Python, FastAPI, LangChain / LlamaIndex, Qdrant / pgvector, embeddings, hybrid search, reranking, OpenAI / Azure OpenAI
Services used
No items found.
The table of content
Overview
About the Project
We are looking for an LLM / RAG Engineer to work on an enterprise knowledge assistant used by internal teams to search policies, product documentation, support articles, technical manuals, and operational knowledge bases. The system focuses on grounded answers, source attribution, document access rules, and measurable response quality.
The stack includes Python 3.11+, FastAPI, LangChain or LlamaIndex where useful, Qdrant or pgvector, hybrid search, embeddings, reranking, document parsing pipelines, and LLM provider integrations. The role includes close communication with international engineering, product, and domain stakeholders, so strong spoken English is essential.
What You Will Do
- Design and implement RAG pipelines for structured and unstructured enterprise documents.
- Work on chunking strategies, metadata extraction, access-aware retrieval, hybrid search, and reranking.
- Integrate vector search with backend APIs and user-facing knowledge assistant workflows.
- Evaluate answer quality, retrieval quality, hallucination risk, latency, and cost.
- Build ingestion pipelines for PDFs, web pages, internal documents, and knowledge base content.
- Implement source citations, confidence indicators, guardrails, and fallback behaviours.
- Collaborate with product and domain teams to define evaluation scenarios and acceptance criteria.
What We Are Looking For
- 3+ years of commercial software engineering experience, with hands-on LLM/RAG project experience.
- Strong Python backend skills and experience with FastAPI or similar frameworks.
- Experience with vector databases such as Qdrant, p