[Hiring] Paid task: fine-tune Qwen2.5-7B-Instruct on our book data and beat our blind-exam score

We are putting the knowledge of one book into the weights of Qwen2.5-7B-Instruct with LoRA and we want someone to beat our current recipe. Paid, 7 days, remote.

Where we are:

  • The book is broken into 2176 units (fact, definition, example, link).
  • After fine-tuning the model started making up answers about things the book does not cover. Refusal examples in the data fixed most of it.
  • Our current model on the blind exam: 2 made-up answers out of 60, 2 refusals out of 15 on practical questions, 28 correct out of 100.
  • An independent examiner reads all answers blind. Thresholds are written down before we see the numbers.

What you get: the 2176 units and our refusal examples (JSON), a passport of our current model (numbers, config, what did not work), a 100-question dev set, and a description of the exam. The test set stays closed.

What you send in 7 days: a LoRA adapter (safetensors), training config and log, your dev numbers, and 10 lines on what you did and why. Any method: SFT, DPO, ORPO, KTO, continued pretraining, changed data. Base must stay Qwen2.5-7B-Instruct, no knowledge from other sources.

How we measure: adapter merged into the base, q8 in ollama, same system prompt as ours, fresh blind set. Three numbers: made-up answers out of 60, refusals on practical questions without the book out of 15, correct answers out of 100.

Pay: $193 + $289. The first part is for any honest run, even if it is worse than ours. The second is if you get at most 2/60 made-up answers, at most 2/15 refusals on practical questions and at least 34/100 correct at the same time. The contract goes through Upwork so both sides have escrow.

To apply, send me a direct message with:

  1. A link to one public artifact of your own fine-tuning of an open LLM (7B or larger): repo with training code, HF model, paper or W&B report under your name, with numbers before and after.
  2. Your plan for this task in 5 sentences.

Not a fit: chatbots, prompt work, RAG without changing weights, API integrations.

Update, 2 October: the first candidate delivered the pilot (open Gutenberg book, LoRA on Qwen2.5-7B, blind judge) in 18 hours on a free Kaggle T4, we re-ran the adapter with our own judge and fresh questions, it held up, and she is on the paid run now. The pilot is the standard entry step for everyone and is written up here: https://dali-pilot.13-140-178-130.sslip.io. Reply in this thread with repo, adapter and the three numbers if you take it.