Skip to main content

Command Palette

Search for a command to run...

Fabric RLM: Why Harness Matters

Updated
•View as Markdown
S

Principal Program Manager, Microsoft Fabric CAT helping users and organizations build scalable, insightful, secure solutions. Blogs, opinions are my own and do not represent my employer.

In my previous blog I introduced Fabric-RLM and an overview of how it works. Let me illustrate with a simple example how giving a model a harness to work with helps.

Question : How many days of the week have the letter d in them?

Of course, the answer is 7. I asked nine small and mid-size models from OpenAI, Mistral, Google and Qwen. Eight of them got it wrong. Then I asked the same nine models the same question through Fabric-RLM. All nine got it right.

The model didn't change but the way model was used changed.

Model Direct Fabric-RLM
openai/gpt-4o-mini 2 ❌ 7 ✅
openai/gpt-4.1-mini 2 ❌ 7 ✅
openai/gpt-4.1-nano 2 ❌ 7 ✅
mistralai/mistral-small-24b 3 ❌ 7 ✅
mistralai/ministral-3b 3 ❌ 7 ✅
mistralai/ministral-8b 4 ❌ 7 ✅
google/gemini-2.5-flash-lite 2 ❌ 7 ✅
qwen/qwen-2.5-7b-instruct 5 ❌ 7 ✅
qwen/qwen3-8b 7 ✅ 7 ✅
Correct 1/9 9/9

Before you say these are old models, here is gpt-5.6-sol :D

Why models get this wrong

LLMs don't read letters, they read tokens. "Monday" is one or two tokens, not six characters. So when you ask whether a word contains a d, the model is guessing from what it has learned about the token, not counting letters. It is a blind spot of how the model sees text, not a lack of knowledge.

How Fabric-RLM got to the answer

Fabric-RLM doesn't give the model any extra knowledge. It gives it a place to work: a Python workspace, and a loop where it can run code, look at the output, and submit an answer only when it's ready.

Here is what gpt-4o-mini did with Fabric-RLM

import os
import dspy
from fabric_rlm import RLM

lm = dspy.LM(
    "openrouter/openai/gpt-4o-mini",
    api_key="sk-<>",
    api_base="https://openrouter.ai/api/v1",
    temperature=0.0,
)

result = RLM.task(
    task="How many days of the week have the letter d in them?",
    outputs={"answer": int},
    lm=lm,
    max_turns=6,
).run()

print(result.payload)          # {'answer': 7}

Fabric-RLM solved it in two turns. First, it wrote down the names of the days in a list and then wrote the Python code to find days with 'd' in them. Same input, same model different harness.

This doesn't mean Fabric-RLM can be used for all tasks. If you ask Fabric-RLM to write a joke, it can't count, filter, compute etc so it's not the right harness this task.

Harness here is the context, tools, guardrails, skills, compute etc the model has to complete the task.