Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RESEARCH · RESEARCH · #1513

THPL: Vision-to-language framework for rainbow trout feeding decisions in RAS

This arXiv preprint (arXiv:2610.02378v1) introduces THPL, a vision-to-language decision-support framework for precision feeding of rainbow trout in Recirculating Aquaculture Systems. The method extracts trajectories with Fishsort to compute an Activity Coefficient (AC), encodes temporal and group dynamics with a Hierarchical Behavior Encoder using Temporal and Set Transformers to produce dual 'physical' and 'soft' tokens, and fine-tunes an LLM with LoRA plus counterfactual multimodal Direct Preference Optimization (mDPO); reported results show strong correlation of AC with expert feeding intensity (Spearman ρ = 0.925, p < 0.001) and decision accuracy improvements from 33.33% (text-only) to 93.33% (with dual-evidence tokens) and to 96.67% after mDPO, along with gains in METEOR and diversity metrics.

KEY POINTS

  1. This arXiv preprint (arXiv:2610.02378v1) introduces THPL, a vision-to-language decision-support framework for precision feeding of rainbow trout in Recirculating Aquaculture Systems.
  2. The method extracts trajectories with Fishsort to compute an Activity Coefficient (AC), encodes temporal and group dynamics with a Hierarchical Behavior Encoder using Temporal and Set Transformers to produce dual 'physical' and 'soft' tokens, and fine-tunes an LLM with LoRA plus counterfactual multimodal Direct Preference Optimization (mDPO); reported results show strong correlation of AC with expert feeding intensity (Spearman ρ = 0.925, p < 0.001) and decision accuracy improvements from 33.33% (text-only) to 93.33% (with dual-evidence tokens) and to 96.67% after mDPO, along with gains in METEOR and diversity metrics.
  3. This work demonstrates a multimodal pipeline that grounds LLM reasoning in continuous spatiotemporal kinematics for operational decision support in precision aquaculture, improving interpretability and decision accuracy.

WHY IT MATTERS

This work demonstrates a multimodal pipeline that grounds LLM reasoning in continuous spatiotemporal kinematics for operational decision support in precision aquaculture, improving interpretability and decision accuracy.

SOURCES & TIMELINE

1