Key Responsibilities
1.1 Design, implement, and operate the in-house LLM Platform backend — API gateway, model routing, request dispatching, and serving abstraction over vLLM, SGLang, TensorRT-LLM, and other inference engines.
1.2 Build OpenAI-compatible and custom LLM API endpoints (chat completions, embeddings, function calling, streaming) for internal product teams, with versioning, backward compatibility, and SDK generation.
...
1.3 Implement platform-level capabilities: authentication and tenant isolation, rate limiting and quotas, usage metering and billing, model fallback and failover, prompt logging and audit trails.
1.4 Integrate with the GPU infrastructure and model registry to route requests to the right model endpoint based on availability, latency, cost, and tenant policy.
1.5 Develop cost-control mechanisms (model routing by tier, token-based metering, caching of prompts and embeddings, request throttling) to keep LLM inference cost-efficient at organization scale.
1.6 Build and maintain the front-end admin console and developer playground (React/TypeScript) for model discovery, API key management, usage dashboards, prompt testing, and request inspection.
1.7 Implement observability for the LLM platform — request tracing, latency and error metrics, GPU utilization dashboards, cost analytics — using Prometheus, Grafana, and distributed tracing.
1.8 Enforce security and governance across the platform — API key management, RBAC, PII/secret redaction in prompts, audit logging, and compliance with internal data policies.
1.9 Collaborate with the GPU training/inference optimization team, model fine-tuning team, and internal product teams to align platform capabilities with evolving model and business needs.
1.10 Research and apply cutting-edge techniques (e.g., semantic caching, speculative decoding integration, multi-model routing, MCP) to optimize platform performance, cost, and developer experience.
Mandatory Qualifications
2.1 3+ years of hands-on experience in backend software engineering, building and operating production APIs and distributed services.
2.2 Deep expertise in Python (or Go) and experience building scalable backend services with frameworks like FastAPI, Flask, Spring Boot, or Go net/http.
2.3 Experience with RESTful API design, streaming APIs (SSE/WebSocket), authentication (OAuth, API keys, JWT), and API versioning.
2.4 Experience with relational databases (PostgreSQL, MySQL), caching (Redis), and message queues (Kafka, RabbitMQ).
2.5 Experience with containerization and orchestration (Docker, Kubernetes) and CI/CD pipelines (Jenkins, GitLab CI, GitHub Actions).
2.6 Working knowledge of modern frontend development with React and TypeScript, including building admin consoles, dashboards, or developer tools.
2.7 Bachelor's or higher degree in Computer Science, Engineering, or related field.
Nice-to-Have Skills
3.1 Experience building or operating LLM serving platforms, AI gateways, or model serving infrastructure (vLLM, SGLang, TensorRT-LLM, Triton, TGI).
3.2 Familiarity with LLM inference concepts — tokenization, KV cache, batching, quantization, streaming completions, function calling, embeddings.
3.3 Experience with OpenAI-compatible API design and LLM SDKs (OpenAI Python/JS SDK, LangChain, LiteLLM).
3.4 Experience with observability and tracing for LLM workloads (Langfuse, OpenTelemetry, prompt logging, evaluation pipelines).
3.5 Familiarity with semantic caching, prompt caching, and cost optimization techniques for LLM APIs at scale.
3.6 Experience with multi-tenant platform design — tenant isolation, quota management, usage metering, and billing.
3.7 Experience with infrastructure-as-code (Terraform, Pulumi) and cloud platforms (AWS, GCP, Azure).
3.8 Contributions to open-source LLM tooling or serving frameworks (vLLM, LiteLLM, LangChain, vLLM ecosystem).
3.9 Experience with GPU workload scheduling on Kubernetes (KubeFlow, Volcano, GPU operator).
3.10 Familiarity with security and compliance for AI platforms (PII redaction, prompt injection defenses, audit logging, RBAC).