코난쌤 블로그

홈전체 글카테고리소개연락처개인정보처리방침

태그: self-distillation

6건의 항목

  • 2026년 8월 19일

    BCSD: 스킬을 잘 쓰는 에이전트를 만드는 양방향 컨텍스트 자기증류

    • agent
    • reinforcement-learning
    • llm
    • skill
    • self-distillation
  • 2026년 8월 11일

    SMRC-SD: 에이전트 자기증류에서 정답지가 현재 상태와 맞지 않을 때

    • agent
    • self-distillation
    • LLM
    • reinforcement-learning
    • multi-turn
    • GRPO
    • harness
    • tool-use
    • loop
    • automation
  • 2026년 8월 09일

    에이전트 RL에서 어떤 액션이 성공에 기여했나 — ADRS의 답

    • agent
    • rl
    • credit-assignment
    • self-distillation
    • grpo
  • 2026년 8월 07일

    AgentOPSD: Agentic RL을 위한 재귀적 자기증류 턴별 크레딧 할당

    • agent
    • reinforcement-learning
    • LLM
    • credit-assignment
    • self-distillation
    • GRPO
    • Bayesian
    • turn-level
    • agentic-RL
    • loop
  • 2026년 8월 05일

    PCSD: 에이전트 RL에서 교사 신호를 토큰별로 신뢰하는 방법

    • agent
    • reinforcement-learning
    • LLM
    • self-distillation
    • agentic-RL
    • credit-assignment
    • GRPO
    • loop
  • 2026년 8월 05일

    TurnSight: 도구 호출 에이전트의 크레딧 할당 문제를 hindsight로 푼다

    • agent
    • reinforcement-learning
    • LLM
    • tool-integrated-reasoning
    • self-distillation
    • hindsight
    • credit-assignment
    • GRPO
    • loop

Created with Quartz v4.5.2 © 2026

  • 소개
  • 연락처
  • 개인정보처리방침
  • 전체 글