# I tested Sonnet 5 in my REAL WORK folder. I was shocked...

https://www.youtube.com/watch?v=XlKVo5wsl8Y

[00:00] And I think it's important to use realworld use cases.
[00:04] And that's why I will show you my personal folder to compare these two models to each other because it's a huge folder with 168 GB and over 140,000 files in there.
[00:16] All crossconnected, interconnected.
[00:19] A team of 50 agents are in this folder.
[00:21] Everything interconnected.
[00:21] And I think that's the right environment that I use on a daily basis for over a year now to see if on a daily work this actually makes sense.
[00:31] So let's dive into this.
[00:33] So I won't use claw desktop for this because that's something that makes in my opinion no sense to use as a professional as a non-coder.
[00:40] It makes no sense to use claw desktop.
[00:42] There you have it.
[00:45] Even that you can use claw code at least use claw code on desktop not co-work to get the most out of these models for the same pricing.
[00:54] The main reason for this is that on claw co-work it's still using just one single agent versus in claw code it's spinning off
[01:02] sub agents and even sub agents and doing the workflows and things like that.
[01:06] So if you really have more complicated work and you don't just want to just have a chatbot talking to you that's what I would recommend.
[01:12] And then obviously what we keep saying on this channel to use a folder that has instructions to guide the the LLM the model whatever you use.
[01:23] This folder just has a team of some agents.
[01:26] These agents have independent agents.mmd notclaw.mmd where they have instructions what they can do.
[01:31] So this means there's an orchestrator Larry who is picking up what I'm asking for and he delegates then the work to these different agents who are specific for specific tasks.
[01:42] That's a demo folder in here.
[01:44] I just show you what my personal folder looks like once you start using this over time.
[01:48] And if you open up here the team, you see I'm having over 50 different agents already.
[01:53] And that's just a process where the team hires new agents if there are new needs.
[01:57] And I'm not diving too deep into this in this video.
[01:59] I have many other videos where I
[02:03] talked about this, but that's about it.
[02:04] So, we will now test how this new set versus oos reacts opening inside the folder to see a comparison how it acts there.
[02:12] So, in order to start this, you just rightclick new terminal at folder and here we are.
[02:17] And there's our terminal.
[02:19] And I cannot repeat enough, no reason to be scared of terminals.
[02:21] They are your best friend when it comes to using AI.
[02:23] You launch Claude and here we are.
[02:25] And they already say, "Meet Sonnet 5. smarter and more efficient for everyday work.
[02:31] Yeah, what is everyday work?
[02:34] Where where does it end?
[02:35] When should you switch to OPOS versus Sonnet?
[02:38] That's the reason why all these companies are wasting so much money on tokens because so many people using AI for primitive work where they could have used even haiku in order to do it and they use models like ous or even fable and burn all these tokens for nothing.
[02:53] And that's what you can do if you have a folder like this.
[02:55] You can actually say that different agents using different models.
[02:59] You can just say it literally
[03:04] Say it.
[03:04] So I can say now PEX for example just uses Sonnet 5 now moving forward.
[03:09] And therefore he won't based a lot of tokens when he's just researching the web.
[03:13] Also Larry as the orchestrator is perfectly fine using just set.
[03:16] I don't need to have OPUS because all he needs to understand is the structure of the folder and where to delegate work and so on.
[03:24] Also, I'm always running everything on Opus just to be sure that he's delegating properly because in my folder with over 160 GBTE of files and 140,000 files in my folder, there's a lot more to overview and I have a lot more complicated tasks for him to do, including coding tasks, even as a non-coder.
[03:44] But that's certainly something that you can do if you know perfectly specific agents that don't need to use a high-end model.
[03:50] In order to change the model, I just go to model.
[03:53] It is already here I switched to set.
[03:55] And what we need to understand if I would now another terminal, this terminal would also run on set.
[03:59] And if I switch the model in one, it will also
[04:05] switch the other model.
[04:07] That's why we need to get out of here and I just type claude model and say sonnet.
[04:12] Now this launches with sonnet.
[04:15] You see on top it used sonnet.
[04:17] And now I launch this in another terminal and I say claude model.
[04:20] Here we are.
[04:23] Opus 4.8. eight sonet 5 and now I can make a proper comparison.
[04:28] otherwise it would just switch one terminal to the dominant one.
[04:30] Now bear in mind both models are looking at this folder and there are basic instructions.
[04:35] So this means claude already knows who he is based on a persona that I gave him.
[04:40] Okay.
[04:43] So go to set and say who are you?
[04:45] Here I say also who are you and I hit enter and I hit enter and let's see because speed is really the thing that I'm hoping for.
[04:50] Okay.
[04:53] And set was a lot faster.
[04:56] But as you can see, OPOS was more comprehensive.
[04:58] So what I want to do next, I'm here in VS Code.
[05:01] Now that's my daily driver where I'm working with my local folder.
[05:03] It is here on the side my
[05:06] Personal local folder.
[05:08] You see by the deliverables.
[05:10] If I go to archive, look at all these things that I worked on in this scaffold alone.
[05:15] So that's the 160 GBTE folder that I've been talking about.
[05:17] Here's my PKM.
[05:19] Here you see in my journal.
[05:21] That goes all the way back to 2017 and let's see what it can tell us about what I worked on I core since 2026.
[05:28] So I will launch on the left OPUS and set here we are oppos 4.8 set 5 and now I say create a summary of what we've delivered for 2026.
[05:42] What are the different updates, what are the feature releases and create a HTML report for me to review.
[05:46] Again I will just copy this over and then let's go.
[05:48] I give even a head start to OPUS and again we are just on high effort here for both.
[05:56] So here he's already saying I'm Larry your team orchestrator.
[05:58] He's identifying himself.
[06:00] He says now because that's my folder and that's very established that it's much more likely he picks up the right things.
[06:04] So let me pull together
[06:07] everything we've shipped for my eye since start 2026.
[06:09] Then have Charter build you the HTML.
[06:12] Charter is a team agent that I have in here.
[06:15] Charter infographic designer.
[06:17] She has she has an avatar and uh she has agent files and so on.
[06:23] So that makes sense because these agents are made for making these reports creating these HTML files.
[06:28] I'm using them on a daily basis and that's the typical workflow that I see here.
[06:34] There's no mention of anything like this for now.
[06:36] See, it creates now fork, which is just a random agent that set created here.
[06:43] And here, what should be in there?
[06:45] Let's just use courses and content, platform, and product.
[06:47] I like the question.
[06:49] How should the report be organized?
[06:50] Feature releases versus updates, monthly timeline.
[06:53] Let's make a monthly timeline.
[06:55] That's it.
[06:57] Here, nothing.
[07:00] No question, no nothing.
[07:00] So, for now, and obviously that's something I need to test more than once to confirm that this is happening, but for now it seems that OPOS is picking up his
[07:09] identity as Larry the orchestrator much better than this guy here.
[07:14] He also launches now general purpose extract there.
[07:17] I could now argue he should use PEX, my research agent, but I usually use PEX to reach out to the web to find information.
[07:24] So, I'm fine with this because this is just data gathering.
[07:26] Nothing specific to look at.
[07:28] But you see again, he's launching a whole swarm of agents while Sonnet just launched one agent doing all the job.
[07:35] He's counting session logs.
[07:37] He didn't ask me about anything.
[07:40] Let's see what the end result will be.
[07:41] I think OP is much more responsive here and telling me what's going on.
[07:45] See, I have what I need for the charter brief and it goes into the guideline 003.
[07:47] So the how the scaffold is structured, we have a team knowledge and in here there are guidelines based all on my specific needs and business needs and here's the guideline design system.
[08:00] So that's what he's saying here and he will look into this to pick up the design system that the team created for us and I could show this in a much
[08:10] more nicer way but I don't care because this is for the agents themselves.
[08:14] But I created this design system.
[08:17] My icor is based on this.
[08:19] All the the social media files, the thumbnails, things like this sits in here.
[08:24] And he perfectly got the point.
[08:26] He will brief Charter.
[08:30] Now to launch Charter as a sub agent using this guideline to create this SOP because we also have SOPs.
[08:36] The whole team works according to as humans would work in a company.
[08:38] That's what my PK is all about and that's what makes them so efficient and the quality is the same.
[08:45] So here we go.
[08:48] SOP 7007 build an infographic.
[08:50] That's what she's looking up now.
[08:52] Not yet.
[08:54] He's preparing obviously.
[08:57] And that's it.
[08:59] And here no mention of any like this.
[09:01] You see 150,000 tokens already.
[09:03] We are beyond that because if you sum this up, we are beyond 150.
[09:06] But the question is always about tokens.
[09:08] If I don't get the quality that I need out of the box and I have a lot of back and forth, I would waste a lot more tokens
[09:10] than investing more tokens up front to get something better out of the box.
[09:15] And you can save so much tokens by having a folder like this.
[09:19] Everything is crystal clear for the model how to work in and what to use and what resources to look up and the guidance and the indexes and all these things.
[09:25] And this also allows to add codecs and Gemini any other models on this folder and they would work in a similar manner.
[09:34] So for example, I could perfectly launch here GLM and you see it launches claude but with GLM 5.2 two in this and that's an open-source model and this what we have here this claw interface in a terminal that's called an harness and it is just providing the LLM an environment to work in and GLM can hijack this so you just you know it's just an easy setup that you can do here and now I could have actually launched the same thing here let's see I will just launch this in parallel just for the sake of it to see how far it will go so just so you see that's a complete different LLM it's
[10:11] claw code.
[10:13] It's just inside the harness of claw code, but it is an LLM open source that you could even run locally on your machine if you have a beefy one.
[10:20] And we will get back to this to see what's going on.
[10:22] Oh, there we go.
[10:24] All right.
[10:24] He got also charter on the on the work now.
[10:27] So you see down here, charter is now also looking into the guidelines.
[10:31] When we look here, we see exactly what charter is doing in order to create this HTML.
[10:35] And that's how you get consistent output for whatever you have by using a scaffold like this that gives you the starting point and then you just keep using it.
[10:44] I'm using mine while this was built over a year that I was using the team in this way and now it's available for everybody free to download and we know from our members how life-changing this thing is when it comes to token usage, the insights they get and most importantly the persistent memory because they keep learning the team.
[11:01] They have in team knowledge.
[11:03] They have session logs.
[11:05] If you go just in 2006, you see just for June.
[11:08] These are all the session logs just for June.
[11:11] So these are
[11:13] all sessions that I ran in June.
[11:16] And they note down what was good, what didn't work, what could have been improved here, April.
[11:20] See, and this is something that's building over time.
[11:24] It's crossconnecting things.
[11:26] The team knows.
[11:28] And as I showed you, the team themselves, they have their own journals.
[11:30] I'm sure if I go into Larry's, here we go.
[11:32] Journal.
[11:33] Look at all these journal entries that he just did for himself working with me.
[11:36] So that's why the differentiation just using one agent for everything versus having a whole team and those team members learn individually for their specific tasks.
[11:47] That's a huge game changer.
[11:49] All right, here we go.
[11:53] OPUS is done.
[11:55] So uh Charter is now grabbing the sign tokens and does the same.
[11:57] So they are uh at the same level.
[11:59] This charter is a bit ahead but again I already get a little report up front where I can see how it's going.
[12:03] That makes it easier also to go backwards and understand what happened inside the chat in my opinion.
[12:07] So to me this is no downside even that you could say well it's a waste of tokens because
[12:13] These are all output tokens.
[12:15] But in the end you know here I see nothing.
[12:17] I just get an end result and I cannot follow the thought process of the LLM properly.
[12:22] Obviously, this could be all adjusted,
[12:24] but thinking that Entropic mentioned that Sonnet uses 35% more tokens now than it did before, this will become very expensive, uh, for sure.
[12:35] And that's where I don't see the reason to not use OPUS.
[12:38] But I'm also pretty sure they will launch Oppus 5 after Sonnet 5, and this will become much more expensive again.
[12:46] So, the overall cost for AI will increase.
[12:50] And with Fable coming back, we also are already in the API era where we would need to pay API in order to get Fable quality.
[12:59] I'm sure I will make a video about this too comparing it if it is really worth it because after Fable disappeared, I kept using OPOS as I did before.
[13:07] And if you look now to GLM, see GLM, that's a fraction of the cost of OPUS.
[13:10] That's why I'm testing this.
[13:12] So
[13:14] even it might be slower, I get the same results.
[13:17] Sometimes I even got better results than OPUS when it comes to design but with a fraction of the cost
[13:23] and this is luckily happening because this drives competition in the AI market and we hopefully see improvement in the pricing but you see even as GLM he recognized I'm Larry the your team orchestrator he is grounding himself he looked into the delivery report ah okay
[13:33] he well he started now looking into this because how this folder works we have here deliverables inbox box where sharing everything that we are working on.
[13:49] So I can work on different things in parallel and then I can just have a folder and I can look into these folders and once they are done they can archive it.
[13:58] So when you go into archive you see there's a lot of things that I worked on in the past and that's still there.
[14:01] If I need a later reference and it's all backed up via GitHub as well.
[14:06] So if anything breaks I can roll things back.
[14:08] But this is how I work every day.
[14:11] So I work on different as you can see.
[14:13] There
[14:15] are a lot of other terminals open that I'm currently working on.
[14:19] So just so you see that I'm actually using claw a lot every day.
[14:25] And here we are. The report is ready.
[14:27] And here the report is ready too.
[14:29] You see here down there 12 minutes 52 seconds.
[14:32] And here 13 minutes 51 seconds.
[14:35] So OPUS was slightly faster delivering this report.
[14:37] And now let's have a look how these look like.
[14:39] And GLM is still working on it because obviously it started a lot later.
[14:44] I will just interrupt him here because well I just made my point already clear that this is working too.
[14:49] You see it started working.
[14:51] It launched these general purpose things.
[14:53] I just want to avoid now that it's overwriting my thing.
[14:54] That's usually not something that I do that I let run three different sessions on the same thing and it's usually completely different things that I do in different sessions.
[15:02] uh but I try to focus on one specific thing so the whole team stays focused on the thing and I are clear on what are the different tabs are that we are working in.
[15:15] So I will just close and terminate now the session and let's see
[15:17] what the output is.
[15:19] So here we have the report from set and here we have the report from OPOS.
[15:25] So let's open this up and here we are on the left we have OPOS and on the right we have Sonnet.
[15:29] These are the two reports and let's see what we delivered.
[15:34] those son actually added out of the box more details up front and I like the looks because well this looks actually better to be honest and if I start scrolling here well we have a legend here and here we have the month and now you guys see also what we delivered in my icor in the past month.
[15:49] so here we shipped new courses we shipped also new lessons in March this was an insane push a lot of things a lot of fixing releases updates April that was an important one because we have our live coaching sessions and those get now completely automatically transcribed and uh chapters created for our inner circle members.
[16:10] So you can directly jump into the time stamps.
[16:14] It's many things. So many things as you can see it just got more over time. Here we in May a lot of
[16:19] things June and it's just listed all the things in a very minimalistic fashion to be honest.
[16:25] And here we have literally just a timeline with the dates where something happened.
[16:32] Ah and it split into different things.
[16:33] Okay, that's interesting.
[16:35] So it made a feature releases list to just focus on the feature releases.
[16:39] And here we have fixes that we shipped in the past few months and in the scope note.
[16:41] Well, now I could argue which one is better.
[16:44] Do it manually.
[16:45] I cannot blame the model for this what they chose.
[16:47] Again, it could be by accident that what model has chosen the other way around.
[16:51] But the question is and that was something I cannot really see here how comparable these two are.
[16:57] But this is something I can do by asking now Claude, there is another report that was created earlier.
[17:04] Can you compare it to the one that you created and find the differences?
[17:08] And then I just paste in the link to this report.
[17:10] And this is enough because it will not only look into this, but also into the deliverables folder.
[17:15] Well, in this case, it is really just this HTML page versus
[17:20] OPUS created an inventory source where there's the whole thing written down as MD file in addition.
[17:27] So I don't care really.
[17:29] But obviously this is again more output tokens that happened in using OPOS.
[17:31] What I also could do, I could even point I don't know if you know this, but you can have every every session that you use to chat with has an individual session ID.
[17:41] And I could now use this session, paste it to OPUS, and say look into the other session and see what this agent did different than what you did.
[17:49] Okay, I'm not doing this because in this case, it makes no sense.
[17:53] But if you're working with several sessions in parallel and something gets messed up or you want to bring over something from another session that you think the other session should know about, that's how you do it.
[18:01] And in order to get session down there, you can just hit slash status line and then write down show uh cla session ID and then it show the session ID here.
[18:15] And what this is every chat is stored locally on your machine.
[18:18] So that's an ID that you can see.
[18:20] cannot access it from outside because it is just local on my machine.
[18:24] So every session that you had the full chat history is stored on your machine and is accessible through these ids and that's what you usually see when you use resume and you want to catch up with another session.
[18:35] These are the sessions based on these session IDs that you can pick up and keep working in.
[18:39] So that's why it asks you do you really want to start the session from scratch or do you want to compact it first because by resuming a session and you launched a new session it will load in the whole conversation again and there's a upfront cost of tokens obviously.
[18:55] Okay, here we are.
[18:57] So what is the difference?
[18:59] It made me a table.
[19:01] I could not say make me an HTML table but I think we offended this.
[19:03] So that's the one from Sonnet and that's the one Oppo created.
[19:05] And here we are.
[19:08] 214 comprehensive 214 items OPOS listed versus 47 that were curated feature releases.
[19:17] Oppos 41 feature releases
[19:21] versus 26 for Sonnet fixes for 84 versus 21.
[19:26] What the heck?
[19:29] So Sonnet didn't even read the whole thing.
[19:31] So I'm wondering if Sonnet actually has a 1 million token window.
[19:34] That's the reason why I'm using OPOS and why I use the max content window.
[19:38] So it really in complex environments I wanted to understand it full.
[19:42] Here content items it didn't it didn't add the category at all versus OPOS did 37 content items that we delivered and here it went through the session logs.
[19:52] That's what they are here for.
[19:54] And here it just went into the release lock, change lock, back fill.
[19:58] So that's why you cannot really trust what AI brings back to you.
[20:01] You need to have an understanding what is your expectations that you would expect to get out of it.
[20:06] And this makes then sense to test some every now and then one model over the other to decide then for yourself what's the model that I really want to use.
[20:14] All right.
[20:16] Many of us are confused me included a bit and that's what we want to break down in this chart that maybe many have seen when it comes
[20:23] to the new set 5 release to see if there
[20:26] is a reason for all this. But the
[20:28] problem, the confusion that many people
[20:30] have is that the Sonnet 5 on this chart
[20:33] is more expensive than OPUS 4.08. So
[20:37] let's just let's break down why. If you
[20:39] look at Sonnet 4.6, these are the
[20:41] different effort levels that these
[20:44] models have. So with a lot more effort,
[20:46] these models can achieve better pass
[20:48] rate for this specific testing. So now
[20:51] the problem is if I look at this testing
[20:53] and I use OPUS with high that's the
[20:56] standard usage that I use we end up
[20:58] about 81 maybe a bit more. Okay the
[21:01] problem now is when I'm using high on
[21:03] set I'm much lower with X high I reach a
[21:07] level which is around 79% that is higher
[21:11] than set 4.6 with the maxed out version.
[21:15] Okay so if we compare now sonnet 4.6 six
[21:18] with Sonnet 5, then Sonnet 5 is more
[21:21] cost effective because I need to spend
[21:24] less than sonet 4 than Sonnet 4.6 and
[21:27] get better results. That makes sense.
[21:29] But where it doesn't make sense anymore
[21:31] is as I said comparing with OPUS 4.8.
[21:34] And then when we crank it up to max
[21:36] effort, we reach the same quality as
[21:39] OPUS, but with a much more expensive
[21:42] cost per token. So you see OPUS on high
[21:46] is the same cost as Sonnet X high but
[21:49] delivers better results. The thing is
[21:51] this is a very specific use case and
[21:54] that's why it's hard to argue about this
[21:56] diagram and only using the model will
[21:59] tell what's the difference and my big
[22:01] hope is it will be speed and our belief
[22:04] is in my eyeore that it's not the model
[22:06] that really makes the good results. It
[22:08] is actually the folder that you're using
[22:10] it in where it has the right
[22:11] instructions and the guidelines and the
[22:13] right agents and the skills in order to
[22:15] get the most out of any model that you
[22:18] use. If they say well set is enough to
[22:20] code and so on and now with these price
[22:22] differences to me it makes no sense at
[22:24] least for OPUS 4.8 right now to change
[22:27] anything to set 5. I will just keep
[22:30] everything on Oppus 4.8 8 and let's see
[22:32] what happens with old pulse 5 when it
[22:34] comes to pricing and all these things
[22:36] that might change but for now it's
[22:38] neither speed nor the output quality
[22:41] that convinces me to use set at all in
[22:45] this environment that I used it here
[22:47] with a professional workspace where I
[22:49] actually get done every day with I hope
[22:51] this gave you some insights and ideas
[22:53] for your own setup and if you want to
[22:55] get this folder as I have here obviously
[22:58] without my personal files but this
[23:00] scaffold you can get it for free inside
[23:01] my icon. Download it, get it started.
[23:03] You have an interface from the get- go.
[23:05] You can use it in Obsidian. You can use
[23:07] it in VS Code and you will see there's a
[23:09] huge difference using this scaffold with
[23:12] your AI versus just using AI out of out
[23:15] of the box. I'm curious, did you already
[23:17] test set 5? Let me know in the comments
[23:19] below. What are your thoughts about it?
[23:20] And are you looking forward to get your
[23:23] hands back on Fable? I surely will do
[23:25] and I will talk about this in the next
[23:27] video. I catch you over there.
