Stepan Tytarenko
cat about.md
about
I’m Stepan, a Senior Applied Machine Learning Engineer based in New York, NY. I build infrastructure and evaluation systems for reasoning and agentic models — sandboxed agent execution at scale, distributed evaluation platforms, and the benchmark harnesses that measure whether any of it actually works.
At Labelbox, I founded and architect a managed agent platform that runs 10,000+ concurrent sandboxed containers across Cloud Batch, GKE, and Modal, and a distributed evaluation-as-a-service platform used by frontier-model customers to run reproducible evaluations with CI/CD, experiment tracking, and real-time dashboards. I unified 12+ reasoning, coding, and agentic benchmark suites — including HLE, AIME, Tau-Bench, SWE-Bench, and Terminal-Bench — under a single harness, and founded a multimodal RL Gym for agents operating across GUI, browser, and CLI environments.
Before that, I built and shipped an LLM-powered semantic product-search system at adMarketplace (Amazon Bedrock, SageMaker, pgvector), and I’ve been a Graduate Researcher at Fordham University working on embedding models, LLM robustness, and interpretable hate-speech detection.
ls -la ~/stack
toolbox
[Python] [PyTorch] [Docker] [PostgreSQL] [Redis] [PySpark] [Databricks] [AWS Bedrock] [SageMaker] [RAG] [LiteLLM]
I focus on benchmark design, agentic evaluation, distributed systems, and sandboxed execution — and I still enjoy building things from scratch, from a from-scratch BERT to a from-scratch PyTorch.
cat contact.txt
connect
Find my full background on the CV page, browse a few projects, or reach out on GitHub or LinkedIn.
