Back to feed
Reddit r/MachineLearning·

Contrastive targeted SFT as a mechinterp method - has anyone mapped causal dependency interactions this way? [D]

Signal
45
Hype
25
In three linesResearcher experiments with iterative targeted SFT combined with mechanistic interpretability on a 31B model. Strategy: contrastive training on specific capability dimensions, then circuit ablation to map causal dependencies between dimensions and optimize future training order.
Read source
Your take?
Fine-tuningReasoningEvals

Summary generated by Claude — human-verified