Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it
8.4 relevance
Score Breakdown
technical depth 9
novelty 8
actionability 8
community 8
strategic 7
personal 10
Scored daily by a customisable AI persona to surface the most relevant engineering leadership news.
Model distillation project is highly technical, novel, and directly actionable for AI/ML engineers.
Summary
The post describes a distillation experiment using DeepSeek V4 Flash as a teacher for GPT-OSS-120B on finance tasks, achieving strong benchmark results (83.61% on FinanceReasoning) and releasing 20B open weights. The key finding is that the teacher model's censorship characteristics did not transfer to the distilled student, suggesting distillation can decouple domain knowledge from alignment behaviors. With no comments yet, the discussion is nascent but the claim is notable for open-source AI development.