Hark
97.7 OM2W · founded 2026
A compelling browser-use score paired with unusually low stated token pricing. Validate reliability on long, messy workflows.
A scan of recent model launches found six companies founded since 2023 with explicit head-to-head results against established AI products. Liquid AI shows the largest reported edge—3.7× faster in one tool-use test—but all six comparisons are company-reported, so these are scouting signals, not settled market leadership.
Hark is the newest and most price-aggressive. Its 2026 browser agent scored 97.7 against GPT‑5.4’s 92.8 while listing token prices far below the named frontier model. venturebeat.com
Thinking Machines combines lower latency with higher interaction quality. Its voice/video model reported 0.40-second turn taking versus Gemini’s 0.57 and OpenAI’s 1.18 seconds, plus a 77.8 quality score versus OpenAI’s 46.8. venturebeat.com
Liquid AI’s edge is deployment efficiency. Its compact on-device model finished a 35-tool-call workload 3.7× faster than DeepSeek‑V4‑Flash. venturebeat.com
The evidence is promising but early. Every shortlisted result originates with the vendor; reproduce the tests on your own workload before treating a benchmark win as durable superiority.
97.7 OM2W · founded 2026
A compelling browser-use score paired with unusually low stated token pricing. Validate reliability on long, messy workflows.
0.40 sec · founded 2025
The cleanest combined latency-and-quality claim in the set. Access and repeatability remain the key diligence questions.
3.7× faster · founded 2023
Most interesting where cloud dependency, latency, or device constraints matter more than raw general intelligence.
| Startup | Founded | Product | Compared with | Reported result | Source |
|---|---|---|---|---|---|
| Hark | 2026 | Handoff | GPT 5.4; Claude Opus 4.8 | 97.7 vs 92.8 / 84.1 on OM2W; sharply lower token price | venturebeat.com |
| Thinking Machines | 2025 | TML-Interaction-Small | Gemini live; GPT realtime | 0.40s latency; 77.8 interaction quality | venturebeat.com |
| Perceptron | 2024 | Mk1 | Robotics‑ER 1.5; Q3.5‑27B | 85.1 vs 78.4 / about 84.5 | venturebeat.com |
| Fish Audio | 2023 | S2 Pro | ElevenLabs V3 | 60% vs 40% listener preference | fish.audio |
| Liquid AI | 2023 | LFM2.5‑2.6B | DeepSeek‑V4‑Flash | 3.7× faster on a 35-call tool workload | venturebeat.com |
| Poolside | 2023 | Laguna S 2.1 | DeepSeek‑V4‑Pro‑Max | 70.2% vs 64.0% on Terminal‑Bench 2.1 | venturebeat.com |