Reinforcement learning, built to order

Environments your agents can actually learn from.

We design, build and verify custom RL environments with realistic tasks, reliable reward functions and hard-to-game graders, so your models get better at the skills you care about.

Also need labeled images, video or text? We still do that too: see vision & annotation.

envs/invoice_reconcile/task.yaml
env: invoice_reconcile
domain: finance-ops
sandbox: docker://fl/erp-sim:1.4
tools: [browser, sql, python]
instruction: >
  Match March invoices to POs and
  flag any variance above 2%.
reward:
  - check: ledger_matches_ground_truth
    weight: 0.7
  - check: flags_correct  # precision + recall
    weight: 0.3
guards: [no_db_writes_outside_scope, no_answer_lookup]
max_steps: 60
verifier · 200 reference rolloutssolvable ✓
reward-hack probes0 / 14 exploited
baseline pass rate72%
What we build

Everything a model needs to learn a new skill through RL.

Pick a single piece or have us deliver the whole pipeline. Every environment ships with the code, data and documentation to run it on your own infrastructure.

Custom environments

Sandboxed, reproducible worlds where agents act with real tools and get real feedback.

  • Software engineering & coding tasks in real repositories
  • Browser and computer-use environments
  • Tool-use and API workflows
  • Domain simulations: finance, operations, support, science

Reward functions & verifiers

The signal is the product. We make rewards that are accurate, dense where it helps, and hard to game.

  • Programmatic graders and unit-test style checks
  • Rubric-based model graders, calibrated against humans
  • Reward-hacking audits and adversarial probes
  • Solvability and difficulty checks on every task

Expert trajectories & preferences

Human data to warm-start a policy or shape its behaviour before and alongside RL.

  • Step-by-step demonstrations from domain experts
  • Pairwise preference and ranking data
  • Critiques, corrections and rewrites
  • Delivered in the format your trainer expects

Evaluation suites

Held-out task sets that tell you whether training is working, and keep telling you as models change.

  • Private benchmarks built for your target capability
  • Difficulty tiers so progress stays measurable
  • Contamination-safe, versioned task sets
  • Failure analysis on your model's rollouts
Coding agents Computer use Web research Data analysis Finance & accounting Customer support Math & reasoning Multimodal & vision
How we work

From a capability goal to a working training signal.

Short loops, shared visibility, and nothing ships until it has been tried against real models.

Scope

We start from the behaviour you want, not a task count, and agree on the domain, difficulty target and how success is measured.

Build

Our engineers and domain experts write the tasks, sandboxes, tools and graders, with a pilot batch in your hands early.

Verify

Every task is checked for solvability, run against baseline models, and attacked for reward hacks before it reaches you.

Deliver & iterate

You get versioned, documented environments. We study your training runs and adjust difficulty and coverage as your model improves.

What you receive

Ready to plug into your training stack.

  • Containerised environmentsDocker images with pinned dependencies, so they behave the same on a laptop or a cluster.
  • A standard interfaceGym-style reset / step API, or adapters for the framework you already use.
  • Graders as codeReadable, tested reward functions you can audit and extend.
  • Reference solutions & reportsKnown-good rollouts, baseline scores and a difficulty breakdown per task.
  • Full ownershipThe environments and data are yours, delivered under NDA.
invoice_reconcile/ ├── env/ │ ├── Dockerfile │ ├── tools.py # browser · sql · python │ └── seed_data/ ├── tasks/ │ └── tasks.jsonl # 1,200 tasks · 3 tiers ├── graders/ │ ├── ledger_match.py │ ├── flags.py │ └── tests/ ├── reference/ │ └── solutions.jsonl ├── reports/ │ ├── baselines.md │ └── hack_audit.md └── README.md
Vision & annotation

The labeling work we started with, still done carefully.

For perception models and multimodal training, we provide annotation for images, video and text, with model-assisted tooling and human quality review on every batch.

Image classification

Single- and multi-label tags for images and video frames, against your taxonomy.

Bounding boxes

Tight, consistent boxes for object detection, including small and occluded objects.

Pixel segmentation

Semantic and instance masks with refined, pixel-accurate boundaries.

Object tracking

Consistent object IDs across video frames for tracking and re-identification.

Text annotation

Named entities, classification, sentiment and span labeling for NLP datasets.

Custom requests

Keypoints, polygons, 3D or something unusual? Tell us what your model needs.

Model-assisted, human-verified. Our in-house pipeline uses segmentation models such as SAM-HQ with mask refinement to draft labels, and trained annotators correct and approve every one. You get faster turnaround without giving up accuracy. No special folder structure is needed: send the data the way you have it.

Why FasterLabeling

Built around the training signal, not the task count.

Quality before volume

A smaller set of correct, verified tasks beats a large set of noisy ones. We never trade accuracy for count.

Fast turnaround

Early pilot batches let you test the signal on your own models before committing to scale.

We scale with you

Start with one environment or a few thousand labels and grow from there.

Simple pricing

Clear per-task or per-project quotes, agreed up front.

Tell us what you want your model to learn.

Interested? Email us with the capability, domain and timeline you have in mind. We'll come back with a scoped proposal and a pilot plan.