We are putting the knowledge of one book into the weights of Qwen2.5-7B-Instruct with LoRA and we want someone to beat our current recipe. Paid, 7 days, remote.
Where we are:
- The book is broken into 2176 units (fact, definition, example, link).
- After fine-tuning the model started making up answers about things the book does not cover. Refusal examples in the data fixed most of it.
- Our current model on the blind exam: 2 made-up answers out of 60, 2 refusals out of 15 on practical questions, 28 correct out of 100.
- An independent examiner reads all answers blind. Thresholds are written down before we see the numbers.
What you get: the 2176 units and our refusal examples (JSON), a passport of our current model (numbers, config, what did not work), a 100-question dev set, and a description of the exam. The test set stays closed.
What you send in 7 days: a LoRA adapter (safetensors), training config and log, your dev numbers, and 10 lines on what you did and why. Any method: SFT, DPO, ORPO, KTO, continued pretraining, changed data. Base must stay Qwen2.5-7B-Instruct, no knowledge from other sources.
How we measure: adapter merged into the base, q8 in ollama, same system prompt as ours, fresh blind set. Three numbers: made-up answers out of 60, refusals on practical questions without the book out of 15, correct answers out of 100.
Pay: $193 + $289. The first part is for any honest run, even if it is worse than ours. The second is if you get at most 2/60 made-up answers, at most 2/15 refusals on practical questions and at least 34/100 correct at the same time. The contract goes through Upwork so both sides have escrow.
To apply, send me a direct message with:
- A link to one public artifact of your own fine-tuning of an open LLM (7B or larger): repo with training code, HF model, paper or W&B report under your name, with numbers before and after.
- Your plan for this task in 5 sentences.
Not a fit: chatbots, prompt work, RAG without changing weights, API integrations.