코난쌤 블로그
Search
검색
다크 모드
라이트 모드
탐색기
홈
전체 글
카테고리
소개
연락처
개인정보처리방침
태그: interpretability
5건의 항목
2026년 9월 01일
Diff Mining — 파인튜닝 로짓 차이만으로 숨은 학습 목적을 잡아내는 방법
interpretability
finetuning
llm
auditing
2026년 8월 29일
도구를 호출할지 말지, 토큰 하나로 조절합니다 — Representation Steering으로 튜닝하는 에이전트 도구 사용
agent
llm
tool-use
interpretability
steering
2026년 8월 14일
Mechanist: AI가 AI의 메커니즘을 스스로 발견하는 에이전트 시스템
agent
interpretability
LLM
mechanistic
safety
automation
harness
loop
2026년 7월 21일
SOPHIA: LLM 추론 루프가 늪에 빠졌을 때 — 숨겨진 활성화 벡터로 탈출시키는 방법
LLM
reasoning
activation-steering
self-loop
inference
agent
interpretability
2026년 5월 11일
Claude의 생각을 텍스트로 읽는다 — Natural Language Autoencoders 인터뷰
ai
interpretability
anthropic
safety