A 24-week, stage-gated plan: five core phases from 90% to 60% confidence, two optional branches, one freeze

Asked:“build a plan, best datasets to use, approaches and more”

Eight chronological work packages over weeks 1–24 test whether a staged latent execution plan — tool identity, argument readiness, dependencies, branch probability, revision — is measurable inside an agent and can safely drive progressively more expensive speculation. The dataset stack was picked from 11 candidate benchmarks: core work runs on BFCL V4, AppWorld and τ²-bench, with ToolSandbox for leakage-proof development and GAIA for final real-world latency validation. Confidence figures are subjective technical-success estimates, not measured probabilities.

The plan: weeks, confidence, kill criteria

Core — must pass its gateOptional branch — droppableFreeze — artifact & paper, no new benchmark experiment
Confidence falls by design along the core funnel — 90 → 82 → 70 → 62 → 60% — each drop bought by a numeric gate: +15 points over trace-only, replication within 3 points on 3/4 model families, zero collateral writes, coverage within 3 points under shift, then ≥20% p50 / ≥15% p95 latency cut with <10% waste. github.com
Phases 6–7 (WebArena at 45%, SWE-bench Verified at 35%) are explicitly droppable: WebArena dies if setup takes over two weeks, and exact shell-command prediction is only a negative control — neither may delay the core paper. ar5iv.labs.arxiv.org
Phase 8 is a freeze, not a benchmark: splits and thresholds lock before final evaluation, and the release is trajectory IDs, extraction hooks, probes, scheduler, latency traces and the harness — WebForge-Bench appears only as reference material. chatpaper.com
Development stays leakage-proof on ToolSandbox’s compact deterministic tools before any headline benchmark is touched; the entry gate demands activations beat action- and trace-only controls by ≥15 points at k=1 with no future-observation leakage. liner.com

Dataset matrix: use, optional, avoid as core

Use as core

    Optional / stretch

      Avoid as core

        Human web traces differ from model-generated tool trajectories; fine as pretraining or offline OOD resources only.

        Approach, baselines and statistics

        Start with linear and matched 1M-parameter MLP probes at every layer and horizon k=1,3,5; then a shared low-rank trunk with heads for no-call, tool family, exact ID, argument readiness, dependencies, branch and revision; compare independent-horizon against autoregressive structured decoders, and escalate to a larger architecture only if the simple probes fail a pre-specified gate. Calibrate with split or domain-conditional conformal prediction for selective action; matched counterfactuals, OOD robustness and activation patching/ablation are required, causal steering is optional. Baseline families: global and conditional frequency; first/second-order Markov; prompt+trace sequence models without activations; verbalized forecasts with and without CoT; linear and matched-size MLP probes; QLoRA; a small separate draft speculator; workflow-pattern speculation; no speculation; eager parallelism; and oracle next-call/dependency/speculation ceilings — always at equal compute and equal dollars. Statistics: task-level splits against trajectory leakage, three or more generation seeds, bootstrap 95% CIs over tasks, paired tests for latency and success; report macro results, worst-domain results, calibration, risk–coverage, lead time and negative cases.

        Per-trajectory data schema (log all of it)

        Work packages

        PhaseWeeksWork package & datasetGo/no-go gateEst. successSource

        Benchmark evidence

        BenchmarkTierScaleOfficial metricsRoleSource

        Research plan from live web sources as of 2026-09-02. Two result sets: 8 chronological work packages (timeline, approach, baselines, metrics, numeric gates, subjective technical-success estimates, links) and 11 candidate benchmarks (role, tier, domain, scale, trajectory properties, metrics, access). Success percentages are subjective estimates, not measured probabilities; dataset facts are version-sensitive — cite exact versions and commit hashes in the implementation. Long approach, baseline and scale texts trimmed for space; full detail sits behind each source link. Venue fork: MLSys if systems gains lead; NeurIPS/ICML/ICLR if latent-plan discovery plus OOD/causal evidence is strongest; AAAI/ACL/EMNLP if the benchmark and agent method dominate.

        This report was generated automatically by Keenable SELECT at a user's request, from publicly available web sources linked herein. Keenable does not review, verify, or endorse its contents and makes no representation as to accuracy, completeness, or timeliness; AI-based extraction may contain errors. Nothing in this report is investment, legal, financial, or other professional advice. All trademarks and referenced content remain the property of their respective owners; no affiliation or endorsement is implied. To report an error, rights concern, or request removal: legal@keenable.ai.

        Keenable SELECTAsk your own question
        Made with Keenable SELECT