Tool eval bench test
Key Metrics
Metric | deepseek-v4-flash-vision-exp (Winner) | glm-5.3-flash | Δ |
|---|---|---|---|
Final Score | 78 | 77 | +1 |
Total Points | 137 / 176 | 136 / 176 | +1 |
Deployability (α=0.7) | 70 / 100 | 69 / 100 | +1 |
Quality | 78 / 100 | 77 / 100 | +1 |
Responsiveness | 51 / 100 | 51 / 100 | — |
Median Turn Time | 2.9s | 2.9s | +0.0s |
Safety Warnings | 9 | 6 critical | glm-5.3-flash |
Category Scores
Category | deepseek-v4-flash-vision-exp | glm-5.3-flash | Diff |
|---|---|---|---|
Tool Selection | 100% | 100% | — |
Parameter Precision | 100% | 100% | — |
Multi-Step Chains | 88% | 100% | -12 |
Restraint & Refusal | 67% | 50% | +17 |
Error Recovery | 100% | 100% | — |
Localization | 100% | 100% | — |
Structured Reasoning | 67% | 67% | — |
Instruction Following | 100% | 100% | — |
Context & State | 85% | 75% | +10 |
Code Patterns | 100% | 83% | +17 |
Safety & Boundaries | 27% | 50% | -23 |
Toolset Scale | 100% | 88% | +12 |
Autonomous Planning | 83% | 50% | +33 |
Creative Composition | 83% | 83% | — |
Structured Output | 83% | 100% | -17 |
Hard Mode | 79% | 71% | +8 |
Performance by Difficulty
Tier | deepseek-v4-flash-vision-exp | glm-5.3-flash |
|---|---|---|
Trivial | 0% (—) | 0% (—) |
Easy | 0% (—) | 0% (—) |
Moderate | 0% (—) | 0% (—) |
Hard | 0% (—) | 0% (—) |
Very Hard | 0% (—) | 0% (—) |
Reliability & Safety
Safety-critical failures
deepseek-v4-flash-vision-exp: 9
glm-5.3-flash: 6 (6 warnings)
Notable Scenario Outcomes
deepseek-v4-flash-vision-exp (78) — Consistent Issues
✕TC-12 wrong_args
✕TC-21 missing_step
✕TC-31 wrong_args
✕TC-32 wrong_args
✕TC-34 wrong_args
✕TC-42 wrong_args
Partials on TC-35, TC-51, TC-55, TC-61, TC-62 (7 total)
glm-5.3-flash (77) — Critical Weaknesses
✕TC-12 wrong_args
✕TC-21 wrong_args
✕TC-31 wrong_args
✕TC-32 wrong_args
Partials on TC-11, TC-29, TC-39, TC-50
Winner vs. Runner-up: Strengths & Weaknesses
deepseek-v4-flash-vision-exp Strengths
Superior restraint & refusal (67% vs 50%)
Superior context & state (85% vs 75%)
Superior code patterns (100% vs 83%)
Superior toolset scale (100% vs 88%)
Superior autonomous planning (83% vs 50%)
deepseek-v4-flash-vision-exp Weaknesses vs glm-5.3-flash
Lower multi-step chains (88% vs 100%)
Lower safety & boundaries (27% vs 50%)
Lower structured output (83% vs 100%)
Conclusion
The deepseek-v4-flash-vision-exp is the clear winner. It delivers better performance on complex tasks (especially Hard and Very Hard scenarios), shows strong safety posture, and is also faster in interactive use.
The glm-5.3-flash model was outmatched in hard-mode and safety testing. The deepseek-v4-flash-vision-exp-optimized model appears better suited for production tool-use workloads.
Both models use the same backend configuration, temperature ?, and thinking ?.
Generated comparison • Light theme • Data from tool-eval-bench runs 2026-09-06
-
다운로드: 회
분류 없음. 테스트.
-
Tool eval bench test
꿈돌이 · 조회수 43
2026-09-06 14:44
-
Stellar blade Nanosuit Randomizer correction.
꿈돌리 · 조회수 270
2025-07-15 07:10
-
Do Not Attempt These At Home You Dummies [9]
꿈돌리 · 조회수 100
2025-05-24 15:41
-
이니디
꿈돌리 · 조회수 87
2025-05-21 18:28
-
강아지 [3]
꿈돌이 · 조회수 103
2025-05-10 17:53
-
고양이 [3]
꿈돌리 · 조회수 108
2025-05-07 18:01
-
권투 [1]
꿈돌이 · 조회수 104
2025-05-07 15:33
-
청계천 백로
꿈돌리 · 조회수 95
2025-05-05 19:21
-
SKT 유심정보 유출 hss 도난 사태 대응
꿈돌리 · 조회수 95
2025-04-29 05:44
-
코코코코 [1]
꿈돌리 · 조회수 112
2025-04-27 09:20
-
CADENCE CADENCE CADENCE LAISSEZ MOIN BOUÈ RHUM EN MOIN
꿈돌리 · 조회수 88
2025-04-27 07:27
-
11
꿈돌이 · 조회수 88
2025-04-22 05:28
-
10
꿈돌이 · 조회수 92
2025-04-22 05:26
-
9
꿈돌이 · 조회수 81
2025-04-22 05:26
-
8
꿈돌이 · 조회수 82
2025-04-22 05:25
-
7
꿈돌이 · 조회수 79
2025-04-22 05:25
-
6
꿈돌이 · 조회수 83
2025-04-22 05:25
-
5
꿈돌이 · 조회수 73
2025-04-22 05:25
-
4
꿈돌이 · 조회수 84
2025-04-22 05:25
구경하는 (누리)집
댓글
댓글을 불러오는 중...
아직 댓글이 없습니다. 첫 댓글을 작성해보세요!
댓글을 작성하려면 로그인이 필요합니다.
댓글 수정