코난쌤 블로그
Search
검색
다크 모드
라이트 모드
탐색기
홈
전체 글
카테고리
소개
연락처
개인정보처리방침
태그: self-distillation
6건의 항목
2026년 8월 19일
BCSD: 스킬을 잘 쓰는 에이전트를 만드는 양방향 컨텍스트 자기증류
agent
reinforcement-learning
llm
skill
self-distillation
2026년 8월 11일
SMRC-SD: 에이전트 자기증류에서 정답지가 현재 상태와 맞지 않을 때
agent
self-distillation
LLM
reinforcement-learning
multi-turn
GRPO
harness
tool-use
loop
automation
2026년 8월 09일
에이전트 RL에서 어떤 액션이 성공에 기여했나 — ADRS의 답
agent
rl
credit-assignment
self-distillation
grpo
2026년 8월 07일
AgentOPSD: Agentic RL을 위한 재귀적 자기증류 턴별 크레딧 할당
agent
reinforcement-learning
LLM
credit-assignment
self-distillation
GRPO
Bayesian
turn-level
agentic-RL
loop
2026년 8월 05일
PCSD: 에이전트 RL에서 교사 신호를 토큰별로 신뢰하는 방법
agent
reinforcement-learning
LLM
self-distillation
agentic-RL
credit-assignment
GRPO
loop
2026년 8월 05일
TurnSight: 도구 호출 에이전트의 크레딧 할당 문제를 hindsight로 푼다
agent
reinforcement-learning
LLM
tool-integrated-reasoning
self-distillation
hindsight
credit-assignment
GRPO
loop