Platform
Solutions
Challenges
Platform
Solutions
Challenges

Research
From experiments to insights: Orq.ai research hub
Filter by theme:
Large Language Models
Generative AI
Prompt Engineering
RAG-as-a-Service

LLM-as-a-judge Alignment with Humans
Ranking evaluation data by how often an LLM judge flips its own answer gives you a review queue that is two to three times better than random sampling, and one blind spot you have to plan around.

PII redaction is cheap for RAG, until your users search for people
Redacting PII from a RAG query costs a few points of recall for semantic search, but knocks the right document out of the top result about 60% of the time for person queries.

Three Red-Teaming Frameworks, One Judge Panel, 5,624 Attacks
What 5,624 adversarial attacks across Evaluatorq, PromptFoo, and DeepTeam tell us about attack quality, judge calibration.

Weak judges, strong panel: an ensemble approach to LLM eval
Your eval pipeline has a judge. That judge has biases. Adding two more judges fixes that, but only if they disagree on the right things.

Copy-Pasting Your Prompt Twice Makes LLMs Smarter
The cheapest accuracy upgrade in LLM engineering still works on GPT-5.4, Claude Sonnet 4.6, and Gemini 3 Pro. Here is what 334 runs taught us.

System Prompt Placement for LLM-as-a-Judge
Every team building an LLM-as-a-judge pipeline faces the same question: does it matter where you put your instructions?

Prompt Optimization for Smaller, Higher-Performing Models
We explored using prompt optimization to make cheaper, smaller language models perform as well as expensive, more capable ones.

Can a 14B Model Match a 100B+ Model? We Fine-Tuned 8
Key takeways from fine-tuning 8+ language models on a text classification task, from tiny 0.6B models to 14B behemoths.
Filter by theme:
Large Language Models
Generative AI
Prompt Engineering
RAG-as-a-Service

LLM-as-a-judge Alignment with Humans
Ranking evaluation data by how often an LLM judge flips its own answer gives you a review queue that is two to three times better than random sampling, and one blind spot you have to plan around.

PII redaction is cheap for RAG, until your users search for people
Redacting PII from a RAG query costs a few points of recall for semantic search, but knocks the right document out of the top result about 60% of the time for person queries.

Three Red-Teaming Frameworks, One Judge Panel, 5,624 Attacks
What 5,624 adversarial attacks across Evaluatorq, PromptFoo, and DeepTeam tell us about attack quality, judge calibration.

Weak judges, strong panel: an ensemble approach to LLM eval
Your eval pipeline has a judge. That judge has biases. Adding two more judges fixes that, but only if they disagree on the right things.

Copy-Pasting Your Prompt Twice Makes LLMs Smarter
The cheapest accuracy upgrade in LLM engineering still works on GPT-5.4, Claude Sonnet 4.6, and Gemini 3 Pro. Here is what 334 runs taught us.

System Prompt Placement for LLM-as-a-Judge
Every team building an LLM-as-a-judge pipeline faces the same question: does it matter where you put your instructions?

Prompt Optimization for Smaller, Higher-Performing Models
We explored using prompt optimization to make cheaper, smaller language models perform as well as expensive, more capable ones.

Can a 14B Model Match a 100B+ Model? We Fine-Tuned 8
Key takeways from fine-tuning 8+ language models on a text classification task, from tiny 0.6B models to 14B behemoths.
Filter by theme:
Large Language Models
Generative AI
Prompt Engineering
RAG-as-a-Service

LLM-as-a-judge Alignment with Humans
Ranking evaluation data by how often an LLM judge flips its own answer gives you a review queue that is two to three times better than random sampling, and one blind spot you have to plan around.

PII redaction is cheap for RAG, until your users search for people
Redacting PII from a RAG query costs a few points of recall for semantic search, but knocks the right document out of the top result about 60% of the time for person queries.

Three Red-Teaming Frameworks, One Judge Panel, 5,624 Attacks
What 5,624 adversarial attacks across Evaluatorq, PromptFoo, and DeepTeam tell us about attack quality, judge calibration.

Weak judges, strong panel: an ensemble approach to LLM eval
Your eval pipeline has a judge. That judge has biases. Adding two more judges fixes that, but only if they disagree on the right things.

Copy-Pasting Your Prompt Twice Makes LLMs Smarter
The cheapest accuracy upgrade in LLM engineering still works on GPT-5.4, Claude Sonnet 4.6, and Gemini 3 Pro. Here is what 334 runs taught us.

System Prompt Placement for LLM-as-a-Judge
Every team building an LLM-as-a-judge pipeline faces the same question: does it matter where you put your instructions?

Prompt Optimization for Smaller, Higher-Performing Models
We explored using prompt optimization to make cheaper, smaller language models perform as well as expensive, more capable ones.

Can a 14B Model Match a 100B+ Model? We Fine-Tuned 8
Key takeways from fine-tuning 8+ language models on a text classification task, from tiny 0.6B models to 14B behemoths.
