TimeThink: Synthetic framework and RLVR to elicit compositional reasoning in timeseries LLMs
The authors present TimeThink, a synthetic data framework that generates deterministic atomic and composite question–answer pairs (with reasoning traces) for timeseries multimodal LLMs. They pair this generator with a reinforcement learning with verifiable rewards (RLVR) training strategy that encourages explicit, compositional temporal reasoning; the paper reports that models trained only on the synthetic data outperform strong baselines on both synthetic and some real-world benchmarks.