Eight chronological work packages over weeks 1–24 test whether a staged latent execution plan — tool identity, argument readiness, dependencies, branch probability, revision — is measurable inside an agent and can safely drive progressively more expensive speculation. The dataset stack was picked from 11 candidate benchmarks: core work runs on BFCL V4, AppWorld and τ²-bench, with ToolSandbox for leakage-proof development and GAIA for final real-world latency validation. Confidence figures are subjective technical-success estimates, not measured probabilities.
Human web traces differ from model-generated tool trajectories; fine as pretraining or offline OOD resources only.
Start with linear and matched 1M-parameter MLP probes at every layer and horizon k=1,3,5; then a shared low-rank trunk with heads for no-call, tool family, exact ID, argument readiness, dependencies, branch and revision; compare independent-horizon against autoregressive structured decoders, and escalate to a larger architecture only if the simple probes fail a pre-specified gate. Calibrate with split or domain-conditional conformal prediction for selective action; matched counterfactuals, OOD robustness and activation patching/ablation are required, causal steering is optional. Baseline families: global and conditional frequency; first/second-order Markov; prompt+trace sequence models without activations; verbalized forecasts with and without CoT; linear and matched-size MLP probes; QLoRA; a small separate draft speculator; workflow-pattern speculation; no speculation; eager parallelism; and oracle next-call/dependency/speculation ceilings — always at equal compute and equal dollars. Statistics: task-level splits against trajectory leakage, three or more generation seeds, bootstrap 95% CIs over tasks, paired tests for latency and success; report macro results, worst-domain results, calibration, risk–coverage, lead time and negative cases.
| Phase | Weeks | Work package & dataset | Go/no-go gate | Est. success | Source |
|---|
| Benchmark | Tier | Scale | Official metrics | Role | Source |
|---|
Research plan from live web sources as of 2026-09-02. Two result sets: 8 chronological work packages (timeline, approach, baselines, metrics, numeric gates, subjective technical-success estimates, links) and 11 candidate benchmarks (role, tier, domain, scale, trajectory properties, metrics, access). Success percentages are subjective estimates, not measured probabilities; dataset facts are version-sensitive — cite exact versions and commit hashes in the implementation. Long approach, baseline and scale texts trimmed for space; full detail sits behind each source link. Venue fork: MLSys if systems gains lead; NeurIPS/ICML/ICLR if latent-plan discovery plus OOD/causal evidence is strongest; AAAI/ACL/EMNLP if the benchmark and agent method dominate.