Tech Meridian ← ENTITY INDEX
PROMY MERIDIAN RU

COMPANY · ENTITY #10911

PPO

Related event timeline, sources and context from the news index.

EVENT TIMELINE

2

COMPANIES · 1 SOURCE · AWS Machine Learning

Amazon SageMaker AI demonstrates multi-turn reinforcement learning (MTRL) fine-tuning for search agents with Qwen3.6-27B

Amazon’s blog post describes using Amazon SageMaker AI’s multi-turn reinforcement learning (MTRL) capability to fine-tune a Qwen3.6-27B model for an LLM-powered search agent. The write-up details MTRL features—serverless execution, modular agent-environment interfaces, asynchronous rollouts, built-in algorithms (PPO, CISPO, IS), trajectory observability via MLflow, and evaluation jobs—and shows an enterprise search use case combining BM25 and vector search tools.

6.0

RESEARCH · 1 SOURCE · arXiv cs.AI

Mechanistic audit of self-discovered RL rule Disco103 shows when learning history helps or hinders

arXiv:2609.35897v1 presents the first causal mechanistic audit of a self-discovered reinforcement-learning update rule (Disco103). By surgically pinning, freezing, and transplanting recurrent states while holding meta-parameters fixed, the paper reports three findings: recurrent history expands usable reward scales (a six‑decade window vs three under zero‑pinning), mismatched history penalties stem from perpetual clamping mitigated if imported state is allowed to evolve, and controlling replay retention can reverse apparent adaptation advantages over DQN under environmental change; results are validated against a second rule (OPEN).

7.0