How scalably can we cheaply fine-tune small models for well-defined tasks? Frontier models are bad at clock reading. I fine-tuned a tiny 450M model from Liquid AI to match GPT-5.6 Sol and beat Fable 5.1 on this task. Here’s my fine-tuned open model called Time Wizard.
Attention: 2 HN points · 0 HN comments. Engagement counts are shown as context. Ansar does not treat popularity as significance.
A launch is the first point at which a product can be evaluated.
Strongest verification behind this event, aged, and discounted by how sure we are it belongs to this company.
How much this moves the ecosystem, independent of your interests.
Share of marketing vocabulary detected in the headline and summary.
Distinct kinds of checkable fact found: amounts, versions, measured changes.
These are Ansar’s estimates, not source claims. The summary above was assembled from source text, not generated.
How scalably can we cheaply fine-tune small models for well-defined tasks? Frontier models are bad at clock reading. I fine-tuned a tiny 450M model from Liquid AI to match GPT-5.6 Sol and beat Fable 5.1 on this task. Here’s my fine-tuned open model called Time Wizard.