← Research questions

Research question · Not an established conclusion

When does human–AI collaboration outperform either alone?

Compare task types, complementary errors, and the design of reliance—not just average gains.

Public database records · This page may be cached; not live monitoring or human validation.

Why ask: collaboration is not complementarity

Adding AI does not establish that a combined system beats either contributor. The comparison must use the same task and evaluation criteria, against the better standalone performer—not only a human baseline.

Scope: complementary performance, correlated errors, appropriate reliance, and task type. The records include experimental and classification settings; they do not represent every form of long-term organizational collaboration.

Existing papers & evidence

The relationships below come from an existing public claim, not title similarity. Summaries and relationship text are HIE record interpretations, not verbatim author statements.

Supports

Generative AI at Work ↗

AI assistance improved productivity in a specific customer-support setting, with heterogeneous effects by worker experience.

Source / review level
Abstract reviewed
Evidence strength (relationship assessment)
0.9
AI assessment confidence
0.91
Human review
No human review
Provenance
codex-research-curation:public-evidence-v1 · codex-research-curation · initial-curation-v1
Original source ↗

HIE editorial interpretation · Not a new empirical finding

Conditions, disagreements & unknowns

The records offer different lenses: task type may shape gains, complementary prediction errors may help, and explanations alone do not guarantee appropriate reliance. These are not interchangeable universal laws.

Still unresolved: do effects transfer across populations, models, and time? What division of work reduces shared errors? Which apparent gains reflect changed evaluation criteria?

Verification references: author-institution pages and original papers. Checking these references does not constitute human scientific review.

Proposed hypothesis

Next: compare delegation, not just AI access

Hypothesis: explicit delegation and calibrated reliance could improve combined performance on some tasks. A proposed comparison uses human-only, AI-only, unrestricted collaboration, and bounded delegation with fixed tasks, scoring, model, and sample. Track errors and costs. No recruitment or execution has occurred.

Support would require beating the better standalone baseline and independent replication. No improvement, gains explained only by extra resources, or offsetting risks and costs would challenge or qualify the hypothesis. Preregistration and authorization must precede execution.

Notes & experiments: no executed results yet. Counterevidence, replication designs, and authorized resources can be proposed; no participation is implied.

Participation & contribution boundaries →