Design and develop scalable backend services and REST APIs using Python, FastAPI/Flask/Django.
Build and integrate Generative AI/LLM-based applications, including RAG, AI agents, embeddings, and tool/function calling.
Work with Azure OpenAI and other LLM providers.
Use AI frameworks such as LangChain, LlamaIndex, LangGraph, or similar.
Implement LiteLLM for LLM gateway, routing, fallback, and usage/cost management.
Implement rate limiting, throttling, caching, retries, and API quotas for scalable services.
Use Langfuse for LLM observability, tracing, prompt management, token/cost tracking, and evaluation.
Design and work with PostgreSQL/MySQL, MongoDB, and Redis.
Deploy and manage applications using Microsoft Azure, Docker, and CI/CD.
Develop unit/integration tests, troubleshoot production issues, and optimize performance and cost.
Participate in system design, code reviews, and technical architecture discussions.
Required
5-7 years of software development experience with strong Python and backend development skills and hands-on experience in Generative AI/LLM applications.
Strong understanding of REST APIs, microservices, RAG, embeddings, vector databases, prompt engineering, AI agents, rate limiting, and distributed systems.
Hands-on experience with Azure OpenAI, LiteLLM, and Langfuse is preferred, along with exposure to LangChain/LlamaIndex/LangGraph or similar AI frameworks.
Experience with Microsoft Azure, Docker, CI/CD, databases, Redis, API security, and cloud-native application development is required.
Strong problem-solving, communication, ownership, and ability to build production-grade AI solutions are essential.
Every tech & IT company hiring across India — with AI match scores — on one live map.
Open the map →