Blog 7 log
guided_open_source_qwen_qwen3.6-35b-a3b.console.log
guided_open_source_qwen_qwen3.6-35b-a3b.console.log / 16.0 KB / 233 lines
SRE-Zero Full Eval Plan
+-----------------------------------------------------------------------------+
| Kind | Baseline | Model | Episodes | Output |
|------+-------------------+-------------------+----------+-------------------|
| llm | guided_open_sour� | qwen/qwen3.6-35b� | 1 | guided_open_sour� |
+-----------------------------------------------------------------------------+
[23:39:43] START run=1/1 baseline=guided_open_source run_all_eval.py:225
model=qwen/qwen3.6-35b-a3b episodes=1
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request failed model=qwen/qwen3.6-35b-a3b attempt=1/6 error=Unexpected message content type: NoneType; finish_reason='length'; native_finish_reason='length'; message_keys=['content', 'refusal', 'role']; has_reasoning=False
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=2/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=2/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request failed model=qwen/qwen3.6-35b-a3b attempt=1/6 error=Unexpected message content type: NoneType; finish_reason='length'; native_finish_reason='length'; message_keys=['content', 'refusal', 'role']; has_reasoning=False
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=2/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=2/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request failed model=qwen/qwen3.6-35b-a3b attempt=1/6 error=Unexpected message content type: NoneType; finish_reason='length'; native_finish_reason='length'; message_keys=['content', 'refusal', 'role']; has_reasoning=False
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=2/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=2/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request failed model=qwen/qwen3.6-35b-a3b attempt=1/6 error=Unexpected message content type: NoneType; finish_reason='length'; native_finish_reason='length'; message_keys=['content', 'refusal', 'role']; has_reasoning=False
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=2/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=2/6
SREZERO_LLM request start model=qwen/qwen3.6-35b-a3b attempt=1/6
SREZERO_LLM throttle sleep model=qwen/qwen3.6-35b-a3b seconds=15.0
SREZERO_LLM request success model=qwen/qwen3.6-35b-a3b attempt=1/6
[00:06:00] END run=1/1 baseline=guided_open_source run_all_eval.py:278
model=qwen/qwen3.6-35b-a3b score=40.258
success=0.273 errors=0
output=D:\SRE-Zero\notes\runs\managed\blog-qwen-
easy-agent-styles-2026-06-13\outputs\guided_open
_source_qwen_qwen3.6-35b-a3b_episodes1.json
full sweep �
1/1 guided_open_source | qwen/qwen3.6-35b-a3b | load_balancer_tls_cert_expi� �
SRE-Zero Baseline Marks
+-----------------------------------------------------------------------------+
| Basel� | Model | Marks | Succe� | Reward | Evide� | Inval� | Steps | Erro� |
|--------+--------+-------+--------+--------+--------+--------+-------+-------|
| guide� | qwen/� | 40.3 | 0.27 | 0.380 | 0.70 | 0.00 | 5.64 | 0 |
+-----------------------------------------------------------------------------+
Wrote records and marks to
D:\SRE-Zero\notes\runs\managed\blog-qwen-easy-agent-styles-2026-06-13\target_su
mmaries\guided_open_source_qwen_qwen3.6-35b-a3b.summary.json
SRE-Zero Marks by Difficulty
+-----------------------------------------------------------------------------+
| | | | | | | Root | Correct |
| Diffic� | Baseli� | Model | Marks | Success | Eviden� | Cause | Fix |
|---------+---------+---------+-------+---------+---------+---------+---------|
| easy | guided� | qwen/q� | 40.3 | 0.27 | 0.70 | 0.64 | 0.64 |
+-----------------------------------------------------------------------------+
Wrote run log to
D:\SRE-Zero\notes\runs\managed\blog-qwen-easy-agent-styles-2026-06-13\logs\guid
ed_open_source_qwen_qwen3.6-35b-a3b.run.log