About The Role
The role is for someone who moves beyond basic prompt engineering and understands how to build reliable, scalable GenAI applications, including advanced RAG systems, multi-agent workflows, and fine-tuning pipelines.
The team works at the intersection of applied research and systems engineering, tackling challenges in latency, hallucination reduction, and context optimization.
Key Responsibilities
• Design and deploy production-grade RAG architectures with advanced chunking, re-ranking, and hybrid search capabilities
• Integrate and manage vector databases such as Pinecone, Qdrant, or pgvector to ensure low-latency semantic retrieval at scale
• Build automated evaluation and guardrail frameworks using LLM-as-a-judge patterns and benchmark datasets
• Fine-tune open-source foundational models using parameter-efficient methods like LoRA and QLoRA on secure datasets
• Optimize inference latency and token usage through quantization, caching, and model distillation techniques
• Collaborate with backend engineers to integrate LLM services into existing microservices and RESTful APIs
What We Are Looking For
• 3-6 years of professional software engineering experience, with a minimum of 2 years dedicated to building LLM or GenAI applications in production
• Strong proficiency in Python and hands-on experience with orchestration frameworks such as LangChain or LlamaIndex
• Deep understanding of embedding models, vector search dynamics, and context window management
• Experience with model hosting, inference optimization, and cloud platforms like AWS, GCP, or Azure
• BS or MS in Computer Science, Artificial Intelligence, or a related technical field
• Bonus: Contributions to open-source AI projects, experience with multi-agent frameworks like AutoGen, or published research in NLP/GenAI