Key Responsibilities1.1 Design, implement, and operate the in-house LLM Platform backend — API gateway, model routing, request dispatching, and serving abstraction over vLLM, SGLang, TensorRT-LLM, and other inference engines.1.2 Build OpenAI-compatible and custom LLM API endpoints (chat completions, embeddings, function calling, streaming) for internal product teams, with versioning, backward compatibility, and SDK generation.1.3 Implement platform-level c
Key Responsibilities1.1 Design, implement, and operate the in-house LLM Platform backend — API gateway, model routing, request dispatching, and serving abstraction over vLLM, SGLang, TensorRT-LLM, and other inference engines.1.2 Build OpenAI-compatible and custom LLM API endpoints (chat completions, embeddings, function calling, streaming) for internal product teams, with versioning, backward compatibility, and SDK generation.1.3 Implement platform-level c