Researchers in China and Singapore have developed an experimental system that translates rainbow trout movement into plain-language feeding instructions for recirculating aquaculture systems. In a controlled pilot, the framework—called THPL—combined video-derived behavior, water-quality measurements and expert feeding rules with a fine-tuned language model. The authors report 96.67% decision accuracy on a 30-segment test set, but the study does not establish that the system improves growth, feed conversion, welfare or farm economics.

The work addresses a practical weakness in intensive fish farming: feeding decisions often depend on fixed schedules or an operator's visual judgment. Those methods can miss rapid changes in appetite, while overfeeding can waste feed and degrade water quality. THPL is intended as a decision-support layer that turns observed behavior into structured recommendations covering feed amount, dispensing interval, feeder state and a short rationale.

The team evaluated the approach at the Yuhang Smart Aquaculture Research Center in Hangzhou. The experiment followed 78 rainbow trout, averaging 65 grams, in a single 1.6-meter-by-1.6-meter tank over 10 days. Fish were fed twice daily at fixed five-second intervals until satiation, while a top-view camera recorded at 20 frames per second. The researchers assembled 299 five-second video segments: 239 from days one through eight for training, 30 from day nine for validation and 30 from day 10 for testing.

Top-down illustration of trout trajectories becoming abstract behavioral evidence tokens.
The pilot derived motion features from overhead video and encoded group behavior as evidence for the decision model. Original editorial artwork.

The first stage used the Fishsort tracking method to estimate individual trajectories from the overhead video. From displacement, velocity and acceleration, the researchers calculated an Activity Coefficient meant to summarize the school's feeding response. That coefficient had a Spearman correlation of 0.925 with expert assessments of feeding intensity, with a reported p-value below 0.001. The result indicates strong agreement in this experiment, though it comes from one tank and one continuous cohort.

THPL does not pass the Activity Coefficient itself to the language model. Instead, a hierarchical behavior encoder converts the tracked motion into two complementary representations: explicit tokens describing physical movement patterns and softer learned tokens intended to retain latent temporal and group-level signals. The encoder uses temporal and set-based transformer components so that the system can represent both how behavior changes over time and how fish move as a group.

Those behavioral tokens are combined with temperature, pH and dissolved-oxygen readings, along with task metadata and a catalog of expert rules. A Llama 3.1 8B model was adapted with low-rank fine-tuning, then further aligned with a counterfactual multimodal preference-training method. The system also applies a deterministic validator to the generated instruction before it can be passed toward feeder control, adding a rule-based check after the language model's response.

Illustration of behavior, water data, and expert rules passing through AI and a safety validation gate.
THPL combines behavioral evidence with environmental readings and expert rules, then checks its structured output before feeder control. Original editorial artwork.

In the paper's ablation tests, a text-only baseline classified feeding decisions correctly in 33.33% of the 30 test segments. Adding both forms of behavioral evidence raised accuracy to 93.33%, and the preference-alignment stage lifted it to 96.67%. The authors also report gains in METEOR and changes in diversity metrics for the generated explanations. With only 30 held-out segments—10 each for weak, medium and strong feeding intensity—the percentages should be read as pilot results rather than a mature performance estimate.

The expert rule catalog maps strong activity to continued feeding at the normal amount and five-second interval, medium activity to a seven-second interval with a 20% reduction, and weak activity to stopping the feeder, waiting 10 seconds and restarting with a 50% probe. These thresholds provide an interpretable operational baseline. The learned model's purpose is broader than reproducing a three-class label: it is designed to express a complete, machine-readable directive and explain which behavioral and environmental evidence supported it.

That distinction also exposes a tradeoff. The researchers acknowledge that if the only goal were to classify feeding intensity, a conventional classifier would be more direct and efficient. Their argument for a language model is that it can serve as a semantic decision interface, combining heterogeneous evidence and rules into a structured action that can be audited by operators or downstream control software.

The largest uncertainties concern generalization and real-world impact. All data came from one cohort, and the test split measures later-day performance within that same cohort rather than independent biological replication. The study did not run a continuous closed-loop deployment or measure feed-conversion ratio, growth, water quality, production cost or welfare outcomes. The authors call for multi-cohort trials, independent replication, closed-loop testing and stronger safeguards for uncertainty, out-of-distribution behavior, human override and fail-safe operation before the approach could support commercial feeding decisions.