GLM-5.2 vs. Claude Opus 4.8: prova de codi 3D
GLM-5.2 i Claude Opus 4.8 creen un cub de Rubik i un joc 3D. Comparem funcionalitat, acabat visual, preu i els límits de la prova.
GLM-5.2 i Claude Opus 4.8 reben exactament dos encàrrecs de programació visual: construir un cub de Rubik interactiu i crear un joc tridimensional de cursa infinita. Tots dos models lliuren prototips funcionals en un sol fitxer HTML, però el vídeo d’AI Stack Engineer troba una diferència clara d’estil. Opus produeix solucions conservadores i netes; GLM afegeix més ambientació, controls i detalls sense que l’usuari els demani.
La prova és interessant com a demostració, no com a benchmark científic. Només hi ha dues peticions, el criteri és principalment visual i no es revisen a fons la mantenibilitat, la seguretat ni el comportament en múltiples execucions. Serveix per veure el caràcter de la primera resposta, però no per declarar un guanyador universal.
1. Per què GLM-5.2 entra en la comparació
El vídeo presenta GLM-5.2 a 00:00 com el nou model insígnia de Z.ai, publicat amb pesos oberts i llicència MIT. La fitxa oficial a Hugging Face especifica 753.000 milions de paràmetres, una arquitectura de barreja d’experts i una finestra de context estable d’un milió de tokens.
Z.ai també introdueix IndexShare, que reutilitza un indexador entre capes d’atenció dispersa. Segons el laboratori, això redueix 2,9 vegades el còmput per token quan el context arriba al milió. El model incorpora nivells d’esforç de raonament i millores en descodificació especulativa per sostenir sessions llargues.
A 01:03, el presentador repassa puntuacions com Terminal Bench, SWE-bench Pro i FrontierSWE. GLM-5.2 s’apropa molt a Opus 4.8 en algunes files i el supera en d’altres, però la taula completa també mostra avantatges amplis d’Opus en SWE-bench Pro, NL2Repo, DeepSWE i ProgramBench. La lectura prudent és que GLM ja competeix a primera línia, no que sigui superior en totes les tasques.
2. Primera prova: un cub de Rubik amb Three.js
El primer encàrrec comença a 03:53. La petició demana un cub de Rubik 3 × 3 amb adhesius de colors, rotació de cares mitjançant arrossegament, càmera orbital, botons per barrejar i reiniciar, i un comptador de moviments. Tot ha de funcionar dins d’un únic fitxer HTML.
Opus 4.8 lliura un cub funcional, estable i fàcil de reconèixer. Els colors són plans, l’estructura és negra i el fons és fosc. Els botons i el comptador responen correctament. És una interpretació literal del requisit, amb una superfície visual reduïda i sense ornaments.
GLM-5.2 resol les mateixes funcions, però a 04:33 el vídeo mostra adhesius arrodonits amb degradats, reflexos, un marc bisellat i una ombra sota el cub. També hi incorpora tecles U, D, R, L, F i B, la notació habitual per girar les cares, i explica que la tecla de majúscules inverteix el sentit.
Aquest detall és la diferència més valuosa de la prova. GLM infereix coneixement del domini i el converteix en una ajuda d’interfície. No és només decoració: ofereix una via addicional per controlar el cub. Opus compleix millor una lectura estricta; GLM intenta anticipar què podria agrair una persona que coneix el trencaclosques.
3. Segona prova: un joc de cursa infinita
El segon prompt arriba a 05:12. Demana un joc Three.js amb terra de quadrícula de neó, un cub que es mou lateralment, obstacles tridimensionals, augment progressiu de velocitat, puntuació, partícules en esquivar per poc i una pantalla de final amb reinici.
La versió d’Opus funciona: hi ha una peça cian, una quadrícula porpra, obstacles roses i un indicador de velocitat. El moviment respon i la partida es pot jugar. El presentador, però, considera que l’escena recorda un tutorial inicial de Three.js perquè no inclou fons, atmosfera ni gaire varietat geomètrica.
A 06:05, GLM presenta una ciutat synthwave, un sol porpra a l’horitzó, una carretera ratllada, parets de neó i obstacles de formes diferents. El jugador és un icosaedre blau envoltat per un anell groc. A més de les fletxes sol·licitades, admet A i D.
Una altra vegada, els dos programes s’executen, però GLM dedica més codi a la primera impressió. Això pot ser ideal per a una maqueta de venda, una demostració o un experiment creatiu. També augmenta la superfície que cal revisar: cada extra pot introduir errors, dependències internes o decisions que el producte no necessita.
4. Prototip visual i codi de producció no són el mateix
El mateix vídeo reconeix el límit a 07:08. Opus tendeix a generar una solució més continguda i fàcil d’inspeccionar. Per a una refactorització, una depuració profunda o codi que s’ha de traspassar a un equip, el presentador encara prefereix aquesta prudència.
GLM, en canvi, guanya el seu criteri en prototipatge ràpid i treball visual. Però una comparació robusta hauria de repetir cada prompt diverses vegades, executar proves automàtiques, comptar errors de consola, revisar accessibilitat, mesurar el rendiment i examinar l’estructura del codi. Tampoc sabem si els dos serveis van aplicar la mateixa configuració d’esforç ni si una nova generació produiria el mateix resultat.
La demostració mesura sobretot “quant posa el model a la pantalla al primer intent”. No mesura el cost total de mantenir el resultat durant mesos. Aquests dos objectius poden premiar comportaments oposats.
5. Cost, context i pesos oberts
L’avantatge estructural de GLM-5.2 apareix a 07:39. El vídeo compara els preus vigents en el moment de gravar i situa l’API de Z.ai molt per sota d’Opus 4.8. Els imports són variables i s’han de comprovar abans de contractar, però el posicionament és inequívoc: GLM vol oferir capacitat de frontera amb un cost d’ús inferior.
Els pesos MIT també permeten desplegaments propis i adaptacions comercials. Això no vol dir que executar 753.000 milions de paràmetres sigui gratuït: la infraestructura necessària és fora de l’abast d’un ordinador domèstic i pot requerir quantització o proveïdors especialitzats.
Per la seva banda, la presentació oficial d’Opus 4.8 destaca millores generals, control de l’esforç, fluxos dinàmics per a tasques grans i un mode ràpid. Anthropic ven una experiència integrada i controlada; Z.ai ofereix més llibertat de desplegament. La tria depèn tant del model com de l’operació que l’envolta.
Conclusions
En aquests dos prompts, GLM-5.2 produeix la demostració més vistosa. Afegeix notació útil al cub, una direcció artística més completa al joc i controls alternatius sense necessitar una segona instrucció. Opus 4.8 compleix tots dos encàrrecs amb una implementació més sòbria i, segons el presentador, més còmoda per revisar.
La prova no demostra que GLM sigui millor programador en general. Sí que mostra una alternativa oberta, competitiva i agressiva en preu que pot encaixar especialment bé en prototips visuals. Per a treball de producció, la decisió hauria d’arribar després de proves repetides sobre el repositori real, amb criteris de correcció i mantenibilitat, no només després d’una captura de pantalla espectacular.
Contrast i context
Fonts consultades
-
01
AI Stack Engineer GLM 5 2 VS Claude Opus 4 8 Side by Side Coding Test is Crazy
-
02
Z.ai a Hugging Face GLM-5.2: model card, arquitectura i benchmarks
- 03
-
04
Anthropic Introducing Claude Opus 4.8
Font de treball
Transcripció amb marques de temps
Consulta la transcripció
-
0:00
, obre el vídeo en una pestanya nova
G, L, M, 5.2, just landed on hugging face with full open weights under an MIT license, and the numbers it's posting are forcing a real conversation. This is Z.A.I.'s new flagship, the Chinese lab formally called G-Poo, and the whole thing was trained on Huawei silicon. No Nvidia chips were involved in any of it. That detail alone shifted some assumptions in the industry. The architecture is a mixture of experts set up. Total parameter count sits around 753 billion, but only about 40 billion stay active per token. That selective routing keeps the inference cost
-
0:40
, obre el vídeo en una pestanya nova
manageable on z.ai side and on yours. They paired it with a real, one million token context window, not the kind that falls apart after 200,000 tokens, plus output up to roughly 131,000 tokens in one shot. two reasoning effort modes ship with it, labeled high and max, so you can dial between speed and depth depending on the job. On terminal bench 2.1, GLM 5.2 scored 81.0, a huge jump from GLM 5.1 at 62.0. On SWE bench pro it landed at 62.1, while GPT 5.5 sits at 58.6 on the same test and GLM 5.1 was at 58.4. On Frontier SWE, which measures long horizon coding work over many steps, GLM 5.2 hits 74.4 percent. GPT 5.5 is at 72.6.
-
1:38
, obre el vídeo en una pestanya nova
Claude Opus 4.8, the model I'm comparing it against today, is at 75.1, so an open weight's Chinese model is sitting within one percentage point of Opus on that benchmark. And it costs roughly 1-6th as much per token through the official API. Index share is the architecture trick worth knowing about. It reuses one lightweight indexer across every four sparse attention layers, and the result is a 2.9 times reduction in per token compute at the full 1 million context. They also upgraded the multi-token prediction layer for speculative decoding, which bumps accepted token length by up to 20% translation, it stays fast when you push it into long sessions. CloudFableFive would be the natural comparison here, but FableFive got pulled. It launched on June 9 as Anthropics mythos class model, the tier above Opus, and got disabled
-
2:35
, obre el vídeo en una pestanya nova
three days later. The US Commerce Department sent Anthropics an export control letter at 521pm Eastern on June 12. The directive barred any foreign national from accessing Fable 5 or Mythos 5, including foreign employees inside and Thropic itself. Anthropic couldn't enforce that filter in real time so they shut both models down for everyone. The trigger was a jailbreak demonstration the administration was worried about for cyber security reasons. Fables Out.
-
3:08
, obre el vídeo en una pestanya nova
Opus 4.8 is what's left from the top tier. KimeK 2.7 code from Moonshot shipped right around the same time with a reported 21 point jump on their internal coding benchmark. Mini Max M3 is posting good numbers too. Quen 3.7 from Ali Baba keeps releasing variants. DeepSeek V4 is in the mix as well, but for the argument happening right now in the developer community, GLM 5.2 and Opus 4.8 are the two, everyone keeps comparing. So let me actually push them. I'm using both models through their official chat interfaces, GLM 5.2 on Z.Ayes Chat, Opus 4.8 on Claude.Aye,
-
3:50
, obre el vídeo en una pestanya nova
same prompts on both sides, two builds. First test is a 3D Rubik's Cube using 3.js. Standard 3 by 3, colored stickers, click and drag to rotate any face, orbit camera with mouse, scramble button, reset button, and a move counter. Easy to describe,
-
4:08
, obre el vídeo en una pestanya nova
harder to build cleanly in a single HTML file. Opus 4.8 returned a working cube. Stickers look clean, rotation feels stable. Scramble works, reset works, move counter ticks up properly. The aesthetic is classic and traditional. Flat colored stickers on a black plastic frame, plain dark background, no extras. It reads as a Rubik's cube, functional and tidy. G-L-M-5.2 took the same prompt and pushed it further. Rossi rounded stickers with a soft gradient on every face, a reflective sheen on top,
-
4:43
, obre el vídeo en una pestanya nova
a thicker beveled frame, and a subtle ground shadow under the cube. It also added a keyboard control strip at the bottom labeling U, D, R, L, F, B. The six standard cube notation moves, plus a hint that holding shift gives counterclockwise turns. That's actual cubernalage baked into the interface. Opus didn't include any of that. Both cubes work. GLMs looks like someone who solved a real cube built the UI. Second test is a 3D endless runner called CubeRunner. 3.js, neon grid floor, player cube sliding left and right with arrow keys.
-
5:22
, obre el vídeo en una pestanya nova
Incoming 3D obstacles, speed ramping up over time, score counter, particle effects on near misses, game over screen with restart. Single-H, TML file, no external assets. Opus 4.8 built a run-able game. A cyan cube on a purple grid, pink, and magenta obstacles, spawning ahead, a small, multiplier countertop right, showing speed. Controls respond, you can dodge, it plays. Visually though, it looks like a 3.js intro tutorial. Bear Grid, plain primitive shapes, no skybox, no background, no atmosphere. GLM 5.2 built something I'd actually screenshot for a portfolio.
-
6:05
, obre el vídeo en una pestanya nova
A synth wave city scene with silhouetted skyscrapers, a giant purple sun setting on the horizon, a striped highway running into the distance, glowing pink and cyan neon walls on either side. The player is a faceted ice blue icosahedron, wrapped in a yellow ring. Obstacles vary, cones, pyramids, twisted knots. A flame effect sits in the middle of the lane As a hazard you have to swore around, the UI text at the bottom even reads drift between walls.
-
6:35
, obre el vídeo en una pestanya nova
AD works too, meaning it added WASD fallback controls without being asked. Both games run. GLM's looks shipable. GLM 5.2 added details the prompt didn't request but a real developer would put in. Cube notation labels on the Rubik's build. WASD as a control fallback on the runner. backdrop instead of empty space. That kind of polish usually shows up on the second or third prompt with most models. Here it showed up on the first try, both times. Opus 4.8 deserves fairness too.
-
7:11
, obre el vídeo en una pestanya nova
Both outputs ran without bugs. Both were stable. Opus tends to produce cleaner, more conservative code, and that's not nothing. For a production refactor, for a deep debugging session. For code I'm handing off to a team, I might still pick Opus because the surface area is smaller and easier to review. But for a one-shot demo for a fast prototype for a first impression piece, GLM 5.2 is putting more on the screen. Opus 4.8 through a cloud consumer plan runs $20 a month and a PI pricing is roughly $5 per million input tokens and $25 per million output.
-
7:50
, obre el vídeo en una pestanya nova
GLM 5.2 on the official hosted API is around $1.40 in and $4.40 out per million tokens. The Z.A.I. subscription tier starts near $12.60 a month. And if you have the GPU capacity, you can pull the weights from hugging face under MIT and run it on your own hardware with zero ongoing cost. For solo devs, indie teams, and anyone in a country where US API billing is friction, that math is a real shift. You don't have to commit to opus. You can run a frontier coding model locally if your hardware is there or rent it for pennies if it isn't. The license is permissive enough for commercial use without restriction.
-
8:34
, obre el vídeo en una pestanya nova
GLM went from 5 to 5.1 to 5.2 in a handful of months. Kimmy shipped K 2.7 code. max pushed M3. Quinn keeps cycling. The open weight side is no longer trailing behind closed frontier models. On Frontier SWE, GLM 5.2 sits within 1% of Opus 4.8 while costing about 82% less per token. That's not a gap I expected to see this year, and it's reshaping how I plan my workflow. Opus 4.8 still has its place in my stack, reasoning heavy debugging, long conversation and refactors, anything where I want the most conservative answer in the room. But for fast prototyping, front and work, anything visual or three-dimensional, GLM5.2 is now my default. The cube runner test alone sold me on that.
-
9:28
, obre el vídeo en una pestanya nova
Alright, so, that's it from the video and I hope you enjoyed it if you did, please like this video and subscribe to the channel and I'll see you in the next video.