Reinforcement learning, built to order
Environments your agents can actually learn from.
We design, build and verify custom RL environments with realistic tasks, reliable reward functions and hard-to-game graders, so your models get better at the skills you care about.
Also need labeled images, video or text? We still do that too: see vision & annotation.
envs/invoice_reconcile/task.yaml
env: invoice_reconcile
domain: finance-ops
sandbox: docker://fl/erp-sim:1.4
tools: [browser, sql, python]
instruction: >
Match March invoices to POs and
flag any variance above 2%.
reward:
- check: ledger_matches_ground_truth
weight: 0.7
- check: flags_correct # precision + recall
weight: 0.3
guards: [no_db_writes_outside_scope, no_answer_lookup]
max_steps: 60
verifier · 200 reference rolloutssolvable ✓
reward-hack probes0 / 14 exploited
baseline pass rate72%