Multi-modal Generation Platform
Image, Video & 3D Inference Backend
35%
lower avg latency
25%
cost cut / 1K requests
3
modalities, one API
Overview
A unified backend fronting image, video and 3D generative models, with async queues for compute-heavy inference.
Architected a unified backend integrating image, video and 3D generative AI models behind one API surface, with async job queues absorbing compute-heavy inference so request threads never block. Average inference latency dropped 35% through provider-level optimisation and intelligent request routing — cheap providers for cheap work, expensive ones only when quality demands it. Response caching and prompt compression delivered a 25% cost reduction per thousand requests, and every routing strategy was validated by A/B test rather than assumption.
Stack
- FastAPI
- Python
- Redis
- Async Queues
- Docker
- AWS
- Year
- 2025
- Focus
- AI/ML, Infrastructure, Backend
- Role
- Architecture & delivery
More work
All projectsThreeZinc AI Platform
Multi-Agent Orchestration System
A multi-agent platform orchestrating Instagram and LinkedIn connectors for automated content generation and campaign analytics.
- LangGraph
- LangSmith
- Node.js
- +4
