Tech Meridian ← ENTITY INDEX
RU

TOPIC · ENTITY #398

reward shaping

Related event timeline, sources and context from the news index.

EVENT TIMELINE

1

RESEARCH · 1 SOURCE · arXiv cs.AI

Planning or Learning: Reliability and Cost in Multi-Asset Maintenance

This arXiv preprint empirically compares planning methods and reinforcement learning (RL) for multi-asset bearing maintenance using run-to-failure data. The authors find planning enforces reliability as a hard constraint and yields zero-failure policies whose total cost is largely insensitive to failure-penalty magnitude, while RL optimizes expected cost and often trades preventive maintenance for occasional failures (yielding lower costs under low-penalty regimes but persistent non-zero failures even when penalties are high); they also test lightweight constraint mechanisms (reward shaping, action masking) and provide a unified benchmark protocol for comparing decision approaches.

6.0