Polymath
Polymath builds the simulated environments AI agents train in. It is aimed at the slowest part of teaching an agent to work: producing the worlds it practises in.
About Polymath
Polymath builds simulated worlds where AI agents learn to operate on their own over long stretches of work, rather than one instruction at a time. The problem it is aimed at is a practical one. Training an agent with reinforcement learning needs an environment for it to act in, and building those environments by hand is slow enough that it, rather than the training itself, is what holds teams up. Polymath makes them easier to produce. The company was founded in 2026 by Dylan Ma and Naren Yenuganti, and went through Y Combinator's Winter 2026 batch. The team came out of UC Berkeley, Hume AI, Plaid and Amazon, and it is based in San Francisco.