The Fine-Tuning Handbook: Free
A free 61-page guide to fine-tuning language models: when to fine-tune, which method to use, how to prepare training data, and how to evaluate the result. Download it with the code FERNFLYBLOG.
A free 61-page guide to fine-tuning language models: when to fine-tune, which method to use, how to prepare training data, and how to evaluate the result. Download it with the code FERNFLYBLOG.
Fernfly now has paid plans. Pro is $49 a month or $490 a year, with no per-message fee and no overage bill, because a model running on your own endpoint costs us nearly nothing per conversation. Here is the whole thing, including what changes on Free and the 80% lifetime discount
Our second webinar, in full. Where the tokens actually go in an agent loop, the nine levers that bring the bill down, and three live runs against a real API showing $28,955 become $843 for the same question and the same answer.
Your brain runs a whole mind on less power than a fridge bulb. Physics permits computing ten thousand to a million times better than we manage. The expensive part was never the thinking, it was the commute.
Training one large model can burn 1,287 megawatt-hours. Physics says the same arithmetic could have cost about eight watt-hours. Where the difference goes, what we tried, and what failed.
The last segment of the agentic AI webinar. A Chart.js dashboard driven by plain English, a tool list we wrote by hand, and the forty lines of ordinary code that sit between the model and the app.
Training GPT-3 consumed about 1,287 MWh; thermodynamics says the same bit-operations could in principle have cost about eight watt-hours. A full accounting of that gap, what closes it, and what we tested that failed.
Part two of the agentic AI webinar: we paste a website address on stage and, ten minutes later, a chat widget is navigating that site and flipping its theme. Every step shown, including the wait.
The opening segment of our agentic AI webinar, in full. What an agent actually is, why agents fail at typing rather than thinking, and why the model inside one should be far smaller than you think.
Paste your website address, get a chatbot that can actually use your site. No API spec, no prompt engineering, no account to set up first.
Most AI chatbots talk. The useful ones do things. Woobert is a ⌘K / Ctrl-K command bar that turns a merchant's plain-English request into the right WooCommerce action, and runs it, on a small specialized Fern model.
Comparing Fern to ChatGPT compares two jobs, not two sizes. Chat rewards scale. Intent→action rewards specialization.
A no-signup playground that runs Fern Bud and GPT-4o on the same intent→action tasks. Same tool call, a fraction of the cost. And no tricks: a miss shows as a miss.
What Qwen-AgentWorld's 'Language World Models' taught us about our tiny tool-calling fleet - including the parts that didn't work.
For the 80%, a frontier model is the wrong tool. Not because it's bad, because it's overkill.
We're opening up the Fernfly blog — product updates, deep dives, and the occasional strong opinion about small models.