Private preview: Sofitra Frontier Finance Index
For frontier AI teams

Training infrastructure for systems that must do real work.

Start with a capability gap. Sofitra finds the operating data behind it, reconstructs the world in which the work happens, and delivers tasks, trajectories and rewards that can move model performance.

The mandate

Static datasets show a model the answer. Real environments require it to understand state, choose tools, recover from mistakes and produce an outcome that survives verification.

We build those environments from the operating histories of real companies, so the difficulty comes from the work itself, not a synthetic proxy.

Capability gaps

Built around what your model cannot yet do.

We work backwards from failure: the economic task, the missing context, the required tools, the decision points and the reward signal.

01 / HORIZON

Long-horizon execution

Multi-step work across documents, databases, spreadsheets and messaging systems.

02 / GROUNDING

Source-grounded judgment

Claims and decisions tied to authoritative evidence rather than plausible text.

03 / RECOVERY

Error detection and repair

Tasks that require finding broken work, diagnosing the cause and recovering safely.

04 / CONSTRAINTS

Hidden-rule compliance

Permissions, agreements, policies and edge-case restrictions that cannot be ignored.

Engagement model

From research question to runnable environment.

A narrow pilot can validate task difficulty and verifier quality before scaling into a full training curriculum.

You define
01

Target capability

The work your model must learn or the benchmark it must survive.

02

Observed failures

Examples of current errors, shortcuts, saturation or reward hacking.

03

Runtime constraints

Models, tools, container standards, latency and security requirements.

04

Training objective

RL, supervised trajectories, evals, reward research or a combination.

Sofitra delivers
01

Runnable environment

Resettable state, tools, data, permissions and task interfaces.

02

Task curriculum

Difficulty bands, hidden holdouts, adversarial cases and refresh plan.

03

Trajectories and repairs

Expert demonstrations, rejected paths, corrections and recovery traces.

04

Verifier and evidence pack

Reward code, rubrics, lineage, methodology and evaluation report.

Delivery anatomy

Everything needed to train, inspect and iterate.

Sofitra environments are model-agnostic and designed to fit the lab’s existing orchestration, evaluation and data-governance stack.

RUNTIMEAPI

Container or API runtime

Deterministic resets, seeded state, controlled tool access, event logs and reproducible execution.

TASKSJSON

Structured task manifests

Initial state, success criteria, hidden constraints, available tools, metadata and split assignment.

TRACESLOG

Action-level trajectories

Tool calls, intermediate artifacts, corrections, approvals and outcome-linked reasoning evidence.

REWARDPY

Composable verifier stack

Outcome checks, process checks, expert rubrics, hidden tests and reward-component diagnostics.

Pilot structure

Prove signal before scaling supply.

A focused engagement tests whether the environment exposes a meaningful capability gap and produces stable reward before a broader programme begins.

WEEK 01

Scope

Define capability, model baseline, tools and acceptance criteria.

WEEK 02

Source

Select the data universe and complete rights and privacy review.

WEEK 03

Build

Reconstruct state, package tools and author the first task family.

WEEK 04

Verify

Implement checks, hidden cases and expert adjudication rules.

WEEK 05

Baseline

Run selected models, analyse failures and calibrate difficulty.

WEEK 06

Deliver

Ship runtime, tasks, trajectories, verifier and evidence pack.

NEXT

Scale

Expand domains, cases, horizons and refreshes from new failures.

Engagement

Bring us the failure worth fixing.

Share a target capability and we will propose the data, environment and verifier architecture around it.