RESEARCH · RESEARCH · #589
Unified empirical evaluation of travel-agent itinerary revision under resource disruptions
This arXiv preprint (arXiv:2609.19654v1) presents a systematic empirical comparison of three approaches for revising travel itineraries after resource disruptions: LLM-Z3 full replanning (with Gemini), IPyHOPPER hierarchical plan repair, and the iTIMO local-revision LLM adapter. Using two TREK-derived benchmarks (500 single-disruption cases and 200 feasible compound-disruption cases), the study evaluates effectiveness, plan stability, and computational cost, finding Gemini+LLM-Z3 best on compound disruptions, IPyHOPPER nearly matching single-disruption success while preserving more accepted commitments, and iTIMO consuming substantially more tokens.
KEY POINTS
- This arXiv preprint (arXiv:2609.19654v1) presents a systematic empirical comparison of three approaches for revising travel itineraries after resource disruptions: LLM-Z3 full replanning (with Gemini), IPyHOPPER hierarchical plan repair, and the iTIMO local-revision LLM adapter.
- Using two TREK-derived benchmarks (500 single-disruption cases and 200 feasible compound-disruption cases), the study evaluates effectiveness, plan stability, and computational cost, finding Gemini+LLM-Z3 best on compound disruptions, IPyHOPPER nearly matching single-disruption success while preserving more accepted commitments, and iTIMO consuming substantially more tokens.
- Quantifies practical trade-offs between full replanning, classical repair, and LLM-based adapters for itinerary recovery, guiding choices that balance feasibility, commitment preservation, and compute cost.
WHY IT MATTERS
Quantifies practical trade-offs between full replanning, classical repair, and LLM-based adapters for itinerary recovery, guiding choices that balance feasibility, commitment preservation, and compute cost.