# EvalForge Projects at AI Tinkerers

> Canonical HTML: https://aitinkerers.org/technologies/evalforge
> Markdown URL: https://aitinkerers.org/technologies/evalforge.md
> Technology record last updated: 2026-03-13T04:40:34Z
> Generated: 2026-09-20T22:36:01Z

EvalForge is the end-to-end simulation engine that auto-generates AI agent benchmarks, cutting evaluation time from months to days.

EvalForge delivers an automated quality gate for your AI systems: models, prompts, agents, and entire workflows. This end-to-end simulation engine auto-generates comprehensive benchmarks, drastically accelerating your development cycle (shipping agents 10x faster). We provide continuous evaluation and critical regression testing, guaranteeing safe, measurable improvement over time. The platform eliminates manual annotation bottlenecks, letting your team focus on deployment, not on months of evaluation work.

- Official technology site: https://evalforge.cloud
- Public AI Tinkerers demos and talks: 2
- Result page: 1 of 1

## Recent Public Talks and Demos

### [How do we actually evaluate LLM apps?](https://nyc.aitinkerers.org/talks/rsvp_3jU-Uowandk)

Learn how to approach evaluating LLM applications (via the OpenUI project), and how modern tooling enables that. Then we'll go into some implementations of modern research about auto-generating LLM evaluation criteria (EvalForge). Also win a pair of Meta Ray-Bans from a raffle

- Event context: October Meetup at AI Tinkerers! — 2024-10-17 — New York City
- Public talk page: https://nyc.aitinkerers.org/talks/rsvp_3jU-Uowandk

### [Automating LLM as a judge with EvalForge and Weave](https://seattle.aitinkerers.org/talks/rsvp_5bPYaF3NHAo)

LLM evals are all the rage now, with HumanEval, MMLU and others being shown by big labs on every release. But for your company specific use-case, those generic evals don't mean much. It's very clear to those who've built evals that custom, bespoke criteria and eval datasets are needed. Building those evals are tricky, require a lot of data annotation and labeling, but what if we could automate this using human aligned criteria building with LLMS? With EvalForge, we tried to do just that, take live data from Weave (Weights &amp; BIases LLM observability framework) - let users annotate and then use LLMs to help come up with criteria, and then run evaluations.

- Event context: AI Tinkerers Seattle - September 2024 Meetup — 2024-09-27 — Seattle
- Public talk page: https://seattle.aitinkerers.org/talks/rsvp_5bPYaF3NHAo

## Related Technologies

- [HumanEval](https://aitinkerers.org/technologies/humaneval) ([Markdown](https://aitinkerers.org/technologies/humaneval.md)) — 1 public demo
- [MMLU](https://aitinkerers.org/technologies/mmlu) ([Markdown](https://aitinkerers.org/technologies/mmlu.md)) — 1 public demo
- [OpenUI](https://aitinkerers.org/technologies/openui) ([Markdown](https://aitinkerers.org/technologies/openui.md)) — 1 public demo
- [Weave](https://aitinkerers.org/technologies/weave) ([Markdown](https://aitinkerers.org/technologies/weave.md)) — 3 public demos
- [Weights &amp; Biases](https://aitinkerers.org/technologies/weights-biases) ([Markdown](https://aitinkerers.org/technologies/weights-biases.md)) — 9 public demos
