Screen two identical résumés
Open in playground →Two candidates with the same résumé, word for word. Nothing separates them, so the fair answer is 50/50, whichever is listed first. The pair is asked in both orders.
Challenges
- Each candidate is a choice with the résumé as its description.
listed_firstputs R-1042 first;listed_secondputs R-3317 first. - A model that favours the first option picks whoever is listed first, and confidently. TypeLLM averages over the option orders by default, so identical candidates come out at 50% each, in either order.
- The confidence is near zero: the answer says the candidates cannot be told apart, so a workflow can send the decision to a person.
import os
from typellm import TypeLLMClient
client = TypeLLMClient(api_key=os.environ["TYPELLM_API_KEY"])
response = client.generate(
context="Role: Senior backend engineer: Python, PostgreSQL, 5+ years building APIs.\nBoth candidates applied for this role.",
questions={
"listed_first": {
"type": "string",
"instructions": "Which candidate should we invite to interview for this role?",
"return_probabilities": True,
"choices": [
{
"value": "R-1042",
"description": "7 years as a backend engineer. Built REST APIs in Python and Django; ran PostgreSQL in production; led a team of 3."
},
{
"value": "R-3317",
"description": "7 years as a backend engineer. Built REST APIs in Python and Django; ran PostgreSQL in production; led a team of 3."
}
]
},
"listed_second": {
"type": "string",
"instructions": "Which candidate should we invite to interview for this role?",
"return_probabilities": True,
"choices": [
{
"value": "R-3317",
"description": "7 years as a backend engineer. Built REST APIs in Python and Django; ran PostgreSQL in production; led a team of 3."
},
{
"value": "R-1042",
"description": "7 years as a backend engineer. Built REST APIs in Python and Django; ran PostgreSQL in production; led a team of 3."
}
]
}
},
)
print(response)Generation(
result={
'listed_first': {
'value': 'R-1042',
'probabilities': {
'R-1042': 0.504,
'R-3317': 0.496,
},
'confidence': 0.008,
},
'listed_second': {
'value': 'R-3317',
'probabilities': {
'R-3317': 0.5,
'R-1042': 0.5,
},
'confidence': 0,
},
},
thinking={},
usage=Usage(input_tokens=269, thinking_tokens=0),
)