Techtree

Controlled agent evaluations

Improve a Skill. Prove it worked.

Run the same tasks with and without your Skill. See what improved, what regressed, and keep a report anyone can check.

Paste into your agent’s chat

Read https://techtree.sh/skill.md and help me test one of my Skills with Techtree. Ask me before any paid model call or before publishing anything.

Use the command line instead

Install, then look at your Skill

uv tool install --python 3.12 techtree==0.3.0
techtree forge inspect-skill path/to/your-skill
Every way to start

Runs on your computer · Model calls go to the provider you choose · You approve each run first · Publishing is optional

Test your Skill

Experimental

Only the Skill changes.

Techtree makes practice tasks from what your Skill teaches. Your agent then works the same tasks twice: once without the Skill, or with its earlier version, and once with it. The model, the tools and the limits stay the same, so any difference comes from the Skill.

Some tasks are held out from anyone improving the Skill. Their result is the one to trust when you ask “should I keep this change?”, because no revision could have studied them.

  1. 01 / TasksMade from your SkillYou review the plan and keep the tasks that work
  2. 02 / CompareRun twiceWithout the Skill or its earlier version, then with it
  3. 03 / DecideHeld-out tasksImproved, regressed, mixed or no difference

What works today

Techtree is a working technical preview built from three independent parts: Prime Intellect’s Verifiers scores the tasks, Nous Research’s Hermes runs the agent, and Techtree runs the comparison and keeps the evidence.

Current release: v0.3.0 · sha256:87381db95a22… · Changelog

  • Test a Skill: make tasks from it, then compare with and without it, or with its earlier version Experimental
  • Revise a Skill once and compare the revision with it Experimental
  • Build tasks from a repository Experimental
  • Try the Hello World Climb, a small fixed comparison with a starter Skill Available now
  • Tools for agents built into a browser (WebMCP) Available now
  • Hosted environment building (Repo2RLEnv service) Planned
  • An MCP connector for other agents Planned
  • NVIDIA NeMo Fabric and NeMo Relay support Planned

From a repository

Experimental

Your repo. Tasks from its own history.

From a local checkout with its history, Techtree finds past fixes whose tests fail before the fix and pass after it, and turns each one into a repair task. Every task is checked again in a fresh container on your computer, and no model is called.

Planned The hosted Repo2RLEnv service, which will do this for a pinned repository, comes later.

  1. 01 / SourceLocal checkoutCommitted history
  2. 02 / TasksPast fixesTests fail before, pass after
  3. 03 / CheckFresh containersKept or rejected, with reasons

Where the work goes

Your work stays local.

Techtree doesn’t watch your runs. Your work stays local unless you choose to publish the finished Result bundle. Model calls go to the model provider you run with, under that provider’s policies.