Stepan Tytarenko

cat about.md

about

I’m Stepan, a Senior Applied Machine Learning Engineer based in New York, NY. I build infrastructure and evaluation systems for reasoning and agentic models — sandboxed agent execution at scale, distributed evaluation platforms, and the benchmark harnesses that measure whether any of it actually works.

At Labelbox, I founded and architect a managed agent platform that runs 10,000+ concurrent sandboxed containers across Cloud Batch, GKE, and Modal, and a distributed evaluation-as-a-service platform used by frontier-model customers to run reproducible evaluations with CI/CD, experiment tracking, and real-time dashboards. I unified 12+ reasoning, coding, and agentic benchmark suites — including HLE, AIME, Tau-Bench, SWE-Bench, and Terminal-Bench — under a single harness, and founded a multimodal RL Gym for agents operating across GUI, browser, and CLI environments.

Before that, I built and shipped an LLM-powered semantic product-search system at adMarketplace (Amazon Bedrock, SageMaker, pgvector), and I’ve been a Graduate Researcher at Fordham University working on embedding models, LLM robustness, and interpretable hate-speech detection.

ls -la ~/stack

toolbox

[Python] [PyTorch] [Docker] [PostgreSQL] [Redis] [PySpark] [Databricks] [AWS Bedrock] [SageMaker] [RAG] [LiteLLM]

I focus on benchmark design, agentic evaluation, distributed systems, and sandboxed execution — and I still enjoy building things from scratch, from a from-scratch BERT to a from-scratch PyTorch.

cat contact.txt

connect

Find my full background on the CV page, browse a few projects, or reach out on GitHub or LinkedIn.