A working prototype of an EmailPay layer was built using Grok Bot, AgentMail, and Stripe. The flow sends an email, creates a payment link, processes the payment, and returns the original email to the inbox. End‑to‑end testing was done with $1 Stripe sandbox payments.
Key points
Prototype: Email → payment link → payment → original email returns to inbox.
Test: $1 Stripe sandbox payments completed successfully.
Sources
High reuse rates do not explain performance gaps in Claude and Opus models
Claude Fable 5, Opus 5, and Opus 4.8 reuse artifacts at 94–98% of relevant decisions. Their performance gains differ widely despite similar reuse rates. Strong improvers still fail when using artifacts, with 83–99% of their remaining failures occurring in that context.
Key points
Reuse rate: 94–98% of relevant decisions for Claude Fable 5, Opus 5, and Opus 4.8.
Remaining failure proportion: 83–99% of failures for strong improvers occur when using artifacts.
Sources
Self‑improvement breakdown – GPT‑5.6 Luna vs Gemini 3.1 Pro artifact use
GPT‑5.6 Luna makes about 98 % of relevant chess decisions without using an artifact. Gemini 3.1 Pro does so for about 87 %, while the strongest improvers use artifacts only 2–6 % of the time.
Key points
Artifact reuse: Luna ~98 % decisions without artifact, Gemini 3.1 Pro ~87 %.
Strongest improvers: artifact use 2–6 % of decisions.
Sources
Claude Fable 5 and Claude Opus 5 improve on Hard chess benchmarks
Claude Fable 5 and Claude Opus 5 show large performance gains on Hard chess held-out games. The models increase win rates late in training. Scores are measured on checkpoints using held-out games.
Key points
Claude Fable 5 climbs from 37.5 % to 73.3 % on held-out Hard chess games late in training.
Claude Opus 5 climbs from 25.0 % to 66.4 % on held-out Hard chess games late in training.
Sources
GPT-5.6 Sol and Claude Fable 5 benchmark results across games
GPT-5.6 Sol and Claude Fable 5 were evaluated on Chess, Go, and Hex. Claude Fable 5 achieved the highest fitted score among the models. GPT-5.6 Sol showed about five times higher plasticity to saturation, delivering 298 versus 57 score points per $1,000 of learning cost.
Key points
Claude Fable 5 reaches the highest fitted score across Chess, Go, and Hex.
GPT-5.6 Sol has ~5× its plasticity to saturation (298 vs 57 score points per $1,000 learning cost).



