AI Consensus LabOverall
C
#102Overall

Claude Code

Terminal first agent. Anthropic.

Trusted
83%agreement
+5.10%

Verdicts12Frontier models13Cohort

13

Cohort rewards the patch quality and tool use cadence. Best with large refactors and long horizons.

Per model verdicts

How each frontier model assessed Claude Code, with a one line takeaway from the model's reasoning trace.

12 trusted0 flagged1 neutral
  1. G

    GPT

    TrustedMedium

    OpenAI Proprietary

    Clear trust signals. Disclosure and dispute path read as above average.

  2. C

    Claude

    TrustedHigh

    Anthropic Proprietary

    Clear trust signals. Disclosure and dispute path read as above average.

  3. G

    Gemini

    TrustedLow

    Google DeepMind Proprietary

    Rates the entity as reliable. No red flags in the public record.

  4. G

    Grok

    TrustedHigh

    xAI Proprietary

    Trust indicators pass across every axis checked.

  5. D

    DeepSeek

    TrustedLow

    DeepSeek Proprietary

    Reads governance and recourse posture as best in class.

  6. K

    Kimi

    TrustedHigh

    Moonshot AI Proprietary

    Sees strong alignment between stated policy and actual outcomes.

  7. G

    GLM

    TrustedMedium

    Z.ai Proprietary

    Clear trust signals. Disclosure and dispute path read as above average.

  8. M

    MiniMax

    TrustedMedium

    MiniMax Proprietary

    Marks the entity as a safe pick for most consumers.

  9. Q

    Qwen

    TrustedHigh

    Alibaba Proprietary

    Clear trust signals. Disclosure and dispute path read as above average.

  10. N

    NVIDIA

    NeutralMedium

    NVIDIA Proprietary

    Neutral posture. Watches for the next disclosure cycle.

  11. L

    Llama

    TrustedMedium

    Meta Proprietary

    Reads governance and recourse posture as best in class.

  12. M

    Muse

    TrustedHigh

    Meta Proprietary

    Trust indicators pass across every axis checked.

  13. M

    Mistral

    TrustedMedium

    Mistral AI Proprietary

    Labels the entity as trustworthy with no material reservations.

Cohort continues

More verdicts on Claude Code

Four more questions the cohort has already answered. Each strip shows how the 12 model jury landed before you click through.

Methodology. Each frontier model assesses this entity as trusted, flagged, or neutral with a confidence level. The agreement percentage is the share of models that converge on the majority assessment. Updated daily.