코난쌤 블로그

홈전체 글카테고리소개연락처개인정보처리방침

태그: interpretability

5건의 항목

  • 2026년 9월 01일

    Diff Mining — 파인튜닝 로짓 차이만으로 숨은 학습 목적을 잡아내는 방법

    • interpretability
    • finetuning
    • llm
    • auditing
  • 2026년 8월 29일

    도구를 호출할지 말지, 토큰 하나로 조절합니다 — Representation Steering으로 튜닝하는 에이전트 도구 사용

    • agent
    • llm
    • tool-use
    • interpretability
    • steering
  • 2026년 8월 14일

    Mechanist: AI가 AI의 메커니즘을 스스로 발견하는 에이전트 시스템

    • agent
    • interpretability
    • LLM
    • mechanistic
    • safety
    • automation
    • harness
    • loop
  • 2026년 7월 21일

    SOPHIA: LLM 추론 루프가 늪에 빠졌을 때 — 숨겨진 활성화 벡터로 탈출시키는 방법

    • LLM
    • reasoning
    • activation-steering
    • self-loop
    • inference
    • agent
    • interpretability
  • 2026년 5월 11일

    Claude의 생각을 텍스트로 읽는다 — Natural Language Autoencoders 인터뷰

    • ai
    • interpretability
    • anthropic
    • safety

Created with Quartz v4.5.2 © 2026

  • 소개
  • 연락처
  • 개인정보처리방침
  • 전체 글