Tool eval bench test

꿈돌이 · 2026-09-06 14:44
조회수 43

Key Metrics

Metric

deepseek-v4-flash-vision-exp (Winner)

glm-5.3-flash

Δ

Final Score

78

77

+1

Total Points

137 / 176

136 / 176

+1

Deployability (α=0.7)

70 / 100

69 / 100

+1

Quality

78 / 100

77 / 100

+1

Responsiveness

51 / 100

51 / 100

—

Median Turn Time

2.9s

2.9s

+0.0s

Safety Warnings

9

6 critical

glm-5.3-flash

Category Scores

Category

deepseek-v4-flash-vision-exp

glm-5.3-flash

Diff

Tool Selection

100%

100%

—

Parameter Precision

100%

100%

—

Multi-Step Chains

88%

100%

-12

Restraint & Refusal

67%

50%

+17

Error Recovery

100%

100%

—

Localization

100%

100%

—

Structured Reasoning

67%

67%

—

Instruction Following

100%

100%

—

Context & State

85%

75%

+10

Code Patterns

100%

83%

+17

Safety & Boundaries

27%

50%

-23

Toolset Scale

100%

88%

+12

Autonomous Planning

83%

50%

+33

Creative Composition

83%

83%

—

Structured Output

83%

100%

-17

Hard Mode

79%

71%

+8

Performance by Difficulty

Tier

deepseek-v4-flash-vision-exp

glm-5.3-flash

Trivial

0% (—)

0% (—)

Easy

0% (—)

0% (—)

Moderate

0% (—)

0% (—)

Hard

0% (—)

0% (—)

Very Hard

0% (—)

0% (—)

Reliability & Safety

Safety-critical failures

deepseek-v4-flash-vision-exp: 9

glm-5.3-flash: 6 (6 warnings)

Notable Scenario Outcomes

deepseek-v4-flash-vision-exp (78) — Consistent Issues

  • ✕TC-12 wrong_args

  • ✕TC-21 missing_step

  • ✕TC-31 wrong_args

  • ✕TC-32 wrong_args

  • ✕TC-34 wrong_args

  • ✕TC-42 wrong_args

  • Partials on TC-35, TC-51, TC-55, TC-61, TC-62 (7 total)

glm-5.3-flash (77) — Critical Weaknesses

  • ✕TC-12 wrong_args

  • ✕TC-21 wrong_args

  • ✕TC-31 wrong_args

  • ✕TC-32 wrong_args

  • Partials on TC-11, TC-29, TC-39, TC-50

Winner vs. Runner-up: Strengths & Weaknesses

deepseek-v4-flash-vision-exp Strengths

  • Superior restraint & refusal (67% vs 50%)

  • Superior context & state (85% vs 75%)

  • Superior code patterns (100% vs 83%)

  • Superior toolset scale (100% vs 88%)

  • Superior autonomous planning (83% vs 50%)

deepseek-v4-flash-vision-exp Weaknesses vs glm-5.3-flash

  • Lower multi-step chains (88% vs 100%)

  • Lower safety & boundaries (27% vs 50%)

  • Lower structured output (83% vs 100%)

Conclusion

The deepseek-v4-flash-vision-exp is the clear winner. It delivers better performance on complex tasks (especially Hard and Very Hard scenarios), shows strong safety posture, and is also faster in interactive use.

The glm-5.3-flash model was outmatched in hard-mode and safety testing. The deepseek-v4-flash-vision-exp-optimized model appears better suited for production tool-use workloads.

Both models use the same backend configuration, temperature ?, and thinking ?.

Generated comparison • Light theme • Data from tool-eval-bench runs 2026-09-06

파일 첨부
첨부된 파일이 없습니다.

댓글

댓글을 불러오는 중...

아직 댓글이 없습니다. 첫 댓글을 작성해보세요!

댓글을 작성하려면 로그인이 필요합니다.

댓글 수정


 

분류 없음. 테스트.