Clip transcript
trickle onto X as well. Prompt thieves. And now, this is something that I historically would never really give a model of this performance level, but it's been competent enough so far that I want to see what we get. All right, so here is our Chronos City result. I think it did it a bit quick. It was like 3 minutes and 45 seconds. Okay, it's simple, but it's actually it's understood the task. We do have moving cars. There's even a 1945 looking pedestrian based off of what they're wearing. The buildings are the density of these windows given the size of the buildings is giving me some level of discomfort that I can't accurately describe why, but the cars are moving on the street. They're just cubes, but All right, let's see if the transition effect is there. There is also sound, and it's got an orbiting effect. I mean, this really it's a bad, but it's not that bad from like a Okay, interesting. Like a RuneScape style sound effect between the two. Much happier, much more colorful. Our folks are all in orange and have yellow top hats. The vehicles are also more vibrant, so the entirety of this scene actually changed to reflect a more vibrant color palette of the '60s. It's going to be dark and neon. Yeah, they It always is. And all models so far have basically done the same aesthetic for 2055. So, there seems to be some common Are they wearing sunglasses? I can't imagine they would be, but actually that does look like sunglasses. Looks like a Next up, we have 2005. Okay. More of blue. Is that a taller singular building in the center that was not there before? No, it was. Okay. Oh, they're going to Oh, all right. Everyone is Oh, the cars have wheels on the bottom of them. They're just under the ground plane. Everyone is bald to such a degree that the sun is reflecting off their heads. Uh, let's let's just check 2025. Very green. Okay, kind of the same thing but a different color. The people are now like fully emissive in terms of their color. And then 2055, let's see if the neon yep, okay. Yeah, that's a pretty common thing that happens. Did the building get taller in the center? I really think it did. So again, really like there's nothing to write home about, but like the core foundational elements are put together competently and this was the first try. There's a nice orbiting movement to them. It changes throughout. There are effects between the scenes, there's sound, and it did some level of like reflection of the specific time period like the vibrant colors of the '60s or the neon colors of the '80s or '50s in the 2000 years. So, pretty pretty not bad. All right, so for the last coding test I want to give this another C++ test, but one that I've given very very rarely to any model, so it's less likely to have made its way in the training data I would think, but it needs to generate a 3D racing game that has like low poly early rally game graphical aesthetic to it using C++ because it really did a surprisingly decent job in the skate game all things considered. And keep in mind models of this size are more few and far between, but really the recent comparison that we have for this is the Poolside Laguna model. And so far, this is definitely outperforming it from coding specific tasks. Though, the Poolside Laguna model was like top tier when it came to role-playing and creative writing, so keep that in mind as well. Having this and then that on a DGX Spark may make for an interesting pairing of like a creative writing and then a coding capable model. You could probably do some pretty creative things entirely on your local device with these two models as a pair. Not running at the same time, but like swapping back and forth. It seemed like the game was running for longer than it had anticipated just doing a test run for, so I stopped it and then gave it a description of the issue as well. And this doesn't seem to want to go away, so All right, it was pretty confident that it was working, but now unfortunately All right, good. It quickly killed the processes, so just in terms of like helping do system stuff at a basic level, it's good at, I guess. All right, supposedly it fixed the black screen issue we were getting. All right, it said it was going to rewrite it completely and it did just try compiling it there and there were no errors, so All right, supposedly, and keep in mind the context length here is now it's at 75% utilization, so this rewrote the entire thing because it was just consistently hitting a black screen. Unfortunately, that is still happening and I'm going to call it at this point because it's the context length is kind of getting up there and it also failed to fix this over a few different iterations. Now, the reason I wanted to do this is because it was doing a good job so far, especially with the skate game, and I wanted to make sure when giving it something that was likely less prevalent in the training data, it would still do a good job. Unfortunately, that didn't happen here, but again, it could also just be a fluke. The model performed well overall throughout the tests. Really just in that one, we failed to get anything decent and it didn't really do a good job of actually improving its result which it did for a lot of these other ones even starting with the browser was I mean when we first looked at this and like nothing was working I was like oh no it's going to be one of those videos and I was fortunately proven wrong the subway station FPS it did a good job at iteratively fixing this I think I gave it like three specific things that it needed to fix and it did do them all at varying levels of success our watch website was acceptable And then the city time thing all those simple basically that I think the combination and take away of what I've noticed here specifically is all of these results are simple but foundationally they actually seem to be somewhat competent and together so it seems like some of the core capabilities here are definitely there and actually competent so yeah it's not going to make like an awesome polished city version of this but as a starting point this is really quite acceptable and again I really think that this paired with the hundred and something billion Laguna model that we just tested recently I think that you would have a awesome creative writing model there and then you would have a decent coding model here so these could be some interesting things for machines like this so I wanted to test this and it does seem like it is hypothetically going to become open weight soon just based off of what they've been posting on X which would make a very very compelling option for those systems I don't know how this would perform at a quantization but that's always something that needs to be tested once it is released and out there in the but as a