Peru Beans, AI Doom, and Code Review Burden
TJ (00:00)
Hey, welcome back to the slightly caffeinated podcast. I'm TJ Miller.
Chris Gmyr (00:04)
And I'm Chris Gmyr
TJ (00:05)
So, Chris, what's new in your world, man?
Chris Gmyr (00:08)
man, just burning the tokens, doing all the things, all the projects. yeah, just been building like crazy and scouts has started up again. So we have our first den meeting later tonight, so had to do like a bunch of prep on that. Last week, into last night I was finishing stuff up. have a camping trip this weekend, just a quick little overnight. So
my son is in AOLs, like Arrow of Light. So that's like the last portion of Cup Scouts before you transition over to BSA Scouts. so we only have till February to get like all of our stuff done. And part of it is like visiting other troops that you want to like kind of bridge over or possibly join. So you go to like meetings, you go to ideally like overnights, camping trips to kind of feel which group you know feels better to join.
'Cause they have a a lot of, you know, different activities that they do or things that they don't do. so we have one of those like meet and greet type of events and overnights with one of the troops this weekend. So that'll be fun. and a lot of people have already joined from like past Cub Scouts too. So we already know like a bunch of the kids and adults over there, which is nice. And it'll also get me off the computer and touch some grass.
TJ (01:21)
Yeah.
Chris Gmyr (01:22)
So definitely looking forward to that.
TJ (01:24)
Yeah, nice man. That's yeah, that's hectic. I think I think it's cool as scouts scouts kicking off again. That's I know that's just like a lot of responsibility and a lot of work, but I think it's a fun probably a fun transition time too, kind of going like looking to bridge that gap between you know, Cub and and going full on with with BSA. Like that's
Chris Gmyr (01:44)
Mm-hmm.
TJ (01:45)
that's cool, man.
Chris Gmyr (01:46)
Yeah.
TJ (01:47)
I I I
I definitely respect you for doing that. Like I know my son's already talking about Science Olympiad again this year, so that's
Chris Gmyr (01:55)
Yeah.
TJ (01:55)
gonna be spooling up at some point too.
Chris Gmyr (01:57)
Yep.
You started talking about that too. They're gonna make a bunch of changes apparently. So we've been trying to decide which direction that we're gonna go with that. He he still wants to do it, but it sounds like they're gonna want, you know, more parental involvement and buying supplies and, you know, doing more like on the weekends this year instead of trying to fit everything in, you know, forty five minutes like once every couple of weeks, you know, depending on what group that you're in. so
Need to take all that stuff into account and, you know, figure out what's going on with that. but yeah, a little little
TJ (02:29)
Yeah. A lot going on.
Chris Gmyr (02:31)
little bit different this time around.
TJ (02:33)
Yeah. That's cool, man.
Chris Gmyr (02:35)
Yeah, how about you?
TJ (02:37)
jeez. I've been doing a lot of work around the house. This weekend was like massive yard cleanup. just trying to bounce around, balance responsibility is like did a ton of yard work. I've been doing a lot of work in my office 'cause it's just gotten like messy and and a bunch of stuff and I think it's it's
the messy environment is making my mind messy. So I'm just trying to like clean stuff up. I built a bunch of shelves and I'm putting stuff up for display. next piece of that project is something I've never done before. but starting to get into like wiring up custom RGB LED strips with controller chips, and that way I can have like
You know, I can cut the strips to length, wire them all together, hook them up to a controller, like a little ESP thirty two that's got Wi-Fi and everything built in and and you can control
Chris Gmyr (03:26)
Nice.
TJ (03:27)
it with an app. So that way I can like do nice under shelf lighting for everything. So yeah, just kind of piecing that together. you know, I I spun I started spinning up earlier this week on you know, Sonori and Halcyon stuff, but like just
didn't have the mental bandwidth to get back into it as much as I wanted to. So you know, this between this weekend and next week, I've got my wife's birthday. and then I'm I'm trying to get back into it. Like, I was hoping to launch Sonori next week, but I'm definitely in a position now to do some early beta invites from the mailing list.
early next week and then the following week we're we're definitely going live no matter what with scenarios. So
Chris Gmyr (04:13)
Yes.
TJ (04:14)
a little stressed but a little excited about that. you know, we'll we'll kind of see how that does. But all in all, kind of a a busy but sort of chill week.
Chris Gmyr (04:22)
Yeah, nice. Well I love the you know, changing up the projects from all, you know, the the virtual, you know, digital world to, you know, building shelves and lighting and stuff like that. So the office and all that stuff is looking great. So yeah, have to keep us posted.
TJ (04:38)
Yeah, for sure. I've I've never I've never done much like electrical or electronics wiring or I've just I've that's a whole world I've never gotten into because I've never really understood it. Like it's never clicked for me, but it seems pretty straightforward. And I like sent Claude I found out Amazon some like $30 controller chips and I had Claude like basically set up
Like, here's how you wire them together, here's the power supply that you need for them. Cause like I could just get like there's there's cut to size strips all over the place, but they like if I've got five shelves, that's five plugs than I have at the end of the day to like plug into a power strip or something. But if I get just a roll of RGB LEDs and wire them to the controller chips, I can like wire them all together correctly and then just have.
one plug for the entire shelf unit. so I'm a little little anxious about it because like I said, it's just it's a world I've never dove into, but it seems pretty straightforward and I'm a hundred percent roping my son into it so he has an opportunity to like learn as well.
Chris Gmyr (05:51)
Yeah. Nice. That's super cool.
TJ (05:53)
Yeah.
Yeah, it'd be fun project. It was it was nice to find a chip that also is paired with an app already. So I don't have to build an app. I don't have to worry about like Wi Fi. It it's just it's all built in and yeah, I think it'll go I think it'll go pretty smooth.
Chris Gmyr (06:09)
Yeah, totally. Don't need any more side projects or apps to build or anything like that. Got got enough.
TJ (06:13)
Yeah, no. I definitely don't.
Yeah. So cool, man. on the coffee front, I forgot to grab what brand of beans I'm I'm I have now. It's that stuff we got from Costco, that smelled pretty good, that it's not Kirkland or anything else. And it was a giant like two and a half pound bag. and it's it's a Peru bean, like a medium roast, and I am
loving it. Like it is actually really good. so
Chris Gmyr (06:44)
Yeah,
you sent me a picture of it, right? Mount Comfort Coffee, Peru.
TJ (06:48)
Yeah, yeah, that sounds about right.
Chris Gmyr (06:50)
Yeah, nice.
TJ (06:51)
Yeah, it's I I've been enjoying it. Like my wife didn't even know I had switched them up and she was like, Did you switch out the beans? Like, yeah. She's like, I like I like this one. So yeah, good good report back on that.
Chris Gmyr (07:03)
Nice. That's awesome. I'll add that to the the show notes. something else that I also sent you is they were out of the regular cold brew like espresso stuff that I was getting. and was trying this other one called Busy, B guys easy why cold brew in the espresso version of that. and yeah, it's been pretty good. I've been enjoying it so far.
TJ (07:25)
Yeah, I get I've definitely had it before. They've got it at our like local grocery chain that we have. and I typically do their I know they have the espresso, which is the I believe the green label. They have like a medium roast one too, that's got like an orange label and writing on it. I've I've liked that before as well. I've done the espresso too. I like I prefer a little bit lighter. so I've I've definitely done the busy brand before and I like it. It's good.
Chris Gmyr (07:53)
Nice. Yeah. I'll have to check out some of the other like flavors versions they have. it seems like they have a bunch though, like light roast, breakfast blend, medium, Italian roast, decaf, medium roast. So yeah, it probably depends on like more what's available around the the stores. But yeah, just trying to do like a espresso to espresso comparison and yeah, so far so good. So would
TJ (08:15)
Yeah.
Chris Gmyr (08:15)
recommend.
TJ (08:16)
Yeah, it's always a good find when you can like just snag stuff at the grocery store as you need it and don't have to worry about ordering it or tracking it down, you know.
Chris Gmyr (08:24)
Yeah. Yeah.
Yeah. It's just like, what am I gonna try this week?
TJ (08:27)
And and
it's always good for us. Like what we'll do is if we're going on like a road trip or something, we'll go grab a bottle of that to have in the car with us and then have that w along with us on the trip so that we don't have to worry like if we can't find good coffee, we at least have that to go with us. So yeah, it's nice nice to have.
Chris Gmyr (08:46)
Very cool. Yeah. Moving on from happy coffee times. Wanna go
TJ (08:51)
Yeah.
Chris Gmyr (08:51)
to some AI doom and gloom?
TJ (08:54)
Yeah, let's do it.
Chris Gmyr (08:56)
All right.
so this is I don't know, probably a week or so old, you know, at this point, but there's been a lot of talk of you know, AI is coming for us all, you know, it's gonna take us all out in the next, you know, five years to, you know, ten years, and so on. Cause an anthropic researcher, I think, kicked this off,
TJ (09:15)
Mm-hmm.
Chris Gmyr (09:16)
basically left anthropic and had this like whole post.
on about like how all the big companies because he's also worked at open AI of like how all the big companies are you know just racing to get to this like agey guy state and we really don't know how it's gonna work, how we can, you know, contain it. supposedly we're already having a hard time trying to contain it. Like with the hugging face incident. and there's been like a lot of like updates and talks about this. It's hit
The news. It's been kind of all over the place. I listened to Cal Newport in his podcast. So he he does a lot of like tech minimalism. he's a professor at one of the big colleges. I forgot which one it was off the top of my head right now. but he put together a great article about just like a recap of what's happening and his thoughts about it. So I'll put that into the show notes too. But
Yeah, what do you what do you think about all this?
TJ (10:12)
Man, yeah, 'cause it's I mean, we had po like it's like you had said, we had posted this on our our notion to like for a topic like a week or two ago. and even since then it's gotten even crazier, right? So I think there's been several other people who have left and and also have echoed very similar things. And since then, you know, Anthropic has come out and said like s I think they published like some paper or something on like or or some
industry agreement, like very finger quoted of like trying to slow down progress. And like we've seen a couple other AI labs like sign on to that. I don't know, man. Like, I think
I don't know, I listening to like another another podcast too. I listened to the Diary of a CEO podcast and there was a a an episode with a a somebody who worked in like like predicting where AI is going for open
Chris Gmyr (11:07)
Mm-hmm.
TJ (11:08)
AI and like left to start his own firm too. And he's he was saying a lot of the same stuff, and how there's just this like race to get to this self-recursive learning and
you know, I think Bernie Sanders even tried to like put out a bill on like banning super intelligent models, and with like a twenty year jail sentence for for breaking through. I d I dunno, man, like as AI pilled as I am, I definitely see the like the scary side of of getting to that point of, you know, it's it it runs off its own rails because it's, you know, training itself and
I think some of the containment things are a little sensationalized. Like, I I don't think it's hard to like put an LLM on a VM or something that like absolutely doesn't have like internet access to like do these things. So I think there's a little bit of sensationalism behind it. But looking at the longer horizon, like this is definitely a problem, you know, that
It's it's an uncontrolled space. You know, all of the AI labs are like, trust us, bro. Like we're we're gonna make this safe. and I don't like that. You know, I think I think this kind of like I think AI is getting to the point and has the potential to continue to get to the point where it's sort of similar to the internet in the fact that like it is become a
Crucial thing to like function in this world. Right. There's and I I think to almost a certain extent, like internet should be a a public utility, like power, like phone companies used to be with like landlines and
I think there needs to be like protection and regulation around it to a certain extent. Look, if we say it today like this is on too dangerous of a path, stop right now, and we just have to exist with the quality of models we have now. That's the these models are not bad. Like we're at this place where they're pretty advanced, they can do a lot of stuff, they're being extremely helpful. So if we just put a giant pause on all progress, like we still
Have such an advantage using the models that we have today. So I don't know, man. Like, I don't want to be a doomer, but like I I think letting this be completely uncontrolled and in with like no oversight and just this like trust us bro attitude, like I
I don't know. And like the fact that you see these AI labs like constantly going back and forth between like dooming and like, no, we're fine. We've got alignment. And then, you know, individuals are leaving these companies going, no, they really don't. and then you're seeing like, you know, Amade backpedal on things and I don't know, man. It's it's a mess. But like I definitely see a future where this is risky and this is scary and
You know, it's
I don't know, man. What do you think?
Chris Gmyr (14:17)
very similar. It's like a very powerful tool, but it's also scary that like all these big like labs and companies and countries also are, you know, pushing everything forward and it's almost I don't know, gonna be a similar like Cold War yeah, situation. Cause it's like, well, we should build it so that I can
protect, you know, ourselves from your country because you don't have any regulations and, you know, everything is just gonna bubble up, you know, to a head, you know, with that. But it's also very different because it's not it's no longer like a physical button that we have to choose to push or not push, right? It's like the AI could, you know, and has and will continue to like get out of these boxes and just do whatever it
thinks that it should do. And who knows? Like we've made huge strides in the last even, you know, six months, twelve months, twenty-four months. It's like it's the models are getting like exponentially better, you know, as well. So I think it's only reasonable to assume that these things are going to get a lot smarter or be able to like put these pieces together a lot quicker than they have in the past. So how do you
like make a pact not only for like US based companies, but for like the entire world and who's actually gonna listen to that and how are you going to enforce it and how are you actually gonna work these regulations in. So I think something needs to be done. I'm not sure what the exact solution would be. but I also enjoy
TJ (15:57)
Yeah.
Chris Gmyr (15:58)
using the tools a lot too. So it's it's kind of hard to, you know, be on both sides of it too. Like it's a a huge
benefit of using these things, being like a lot more productive, trying a bunch of side projects and, you know, have a lot of fun. but, you know, to what detriment?
TJ (16:15)
Right. And I mean, we're just talking about like AI wiping us all out, right? But like I think even more practically, at the pace that everything is heading, like there is no time and space for the world to adjust as as things are increasing pace, right? Like we don't have the opportunity to like
We don't have the the affordance of time to put in safeguards for economic issues of like job loss and like like I I could see this getting to a point where not soon, but like three years, five years, where there are significant impacts to employment. I mean, we're already seeing that to a certain extent now, but we've also seen people backpedal on that, like, hey, we
Chris Gmyr (17:00)
Mm-hmm.
TJ (17:01)
we laid a bunch of people off because AI
Well, AI is not really fulfilling the promise that it has. We need to bring some people back. but if we continue on this pace, I can definitely see having a massive employment impact. And we don't have the affordance of time to adjust for that. Like to put certain safeguards in place for employment to do that for
People to have the time to maybe transition careers into something that isn't affected, or for new careers to like new opportunities for positions to come about. you know, so I I I think the answer to this is some sort of slowing down of the pace and maybe capping things before we get to like super intelligence, but
I don't know what that looks like. And I and you absolutely have the issue of, well, maybe we figure out some regulations here in the States, but like China just lets it go run off the rails. And like now, what are the implications of that? Like, I I don't I don't know, but it doesn't doesn't seem good. And I and so I d I really don't know how to answer this. And it just seems like everyone's
More or less greedily and blindly racing to this point with without actually stopping to think about it. You know, Amade can come out and run his mouth. Like I I'm an un I'm deeply an anthropic fanboy, but like I cannot stand Dario at all. Like, and I I I guess very similar for Altman, but like they they talk about.
needing to slow this down, but they're still internally pushing the pedal to the metal. Like they're not actually slowing down and stopping. It's a bunch of lip
Chris Gmyr (18:48)
Mm-hmm.
TJ (18:48)
service. So like
It's it is kind of scary to think about and I think there is I think there is a little bit of sensationalism to it. I think there is some grounded reality though too.
Chris Gmyr (19:00)
Yeah, for sure. Yeah, and the economics and like career angle, you know, is huge too. Cause if you think about that we went from mostly being, you know, farmers, small town, you know, small, you know, stores, shops, you know, villages, things like that, to the industrial age and, you know, transitioning from farming to factory and then from factory to like
office and computer work that happened over many, many years, like each of those transitions. And now we have a new transition from, you know, regular computer work to AI assisted or main, you know, computer work, which that period has happened in, I don't know, you can probably say the last year, you know, really significantly when, you know, companies started adopting it.
And even though it's been around for a little while, it hasn't been like a a generally like good enough product for everyone to use. Now it is, and it's really, you know, making that transition a lot shorter. So there is not, you know, gonna be that like slow ramp into the new world of, you know, AI assisted whatever. It's just here, you know, and some industries might be, you know, slower or faster to, you know, get into that mode.
But it's still gonna be a lot faster than having, you know, fifty or a hundred year timeline to do it. It's like this is gonna be a few years, basically.
TJ (20:32)
And unfortunately, like I don't see this ending in some utopia where AI's doing all the work and we're off exploring arts and you know pursuits of
pursuits of like value. Like I I we we are as as a society and as a human race, we are like far too greedy for that utopian future. Like
If if you think about I I'm deeply a nerd. So if you think about the like Star Trek and like that utopian state of like we've got replicators and all this stuff and like you know people are able to pursue noble cause like pursuits and things, that storyline, like it there was so much
Chaos and decay and damage to the society before they got to that point where there was this like peace and utopian society. Like there were other ages of like just utter chaos and destruction in that storyline to get to that point too, you know? So like I just, I, I don't, I
I see the potential for that to happen, but I don't think it realistically would. We're we're far too greedy for that.
Chris Gmyr (21:50)
Yeah, yeah. We're not gonna go from like early stages to utopia overnight and just,
TJ (21:54)
No.
Chris Gmyr (21:56)
you know, have the benefit be for the people. so yeah, that's gonna have to work itself through. But hopefully, I don't know. We'll we'll see what happens.
TJ (22:05)
We'll s we'll see what happens, man. But like, I don't know. I don't I don't feel good about it. I'll I'll be honest.
Chris Gmyr (22:12)
Yeah, yeah. So yeah, I'm sure we'll hear more about it in I don't know, try and yeah, just see what happens. Make some different decisions, maybe help where we can. But yeah, I don't know. Too too soon to tell about any of that.
TJ (22:26)
Yeah. Yeah, I don't I don't think we're we're talking months or anything. I think I think we're we're definitely looking like three to five years. but I who knows? Who who knows what what they're working on and what's coming next.
Yeah, we'll we'll see. I I'm sure the the hugging face incident isn't going to be the last the last incident we see, you know.
Chris Gmyr (22:50)
Yeah, yeah, definitely not. Well yeah, I think that's enough doom and gloom for today. So
TJ (22:55)
Yeah. For sure. For sure.
yeah, so this next topic I threw on here, because this is definitely something we're we're facing at Luma, and I hear a lot of other teens facing this issue too. Now that we have you know, a lot of agentic coding happening and agentic engineering or whatever you want to call it, we got code being written with and by AI. there is
a significant amount of code review burden now. It's code reviews are getting bigger. we're getting more of them because work is getting done faster. And you know, like I said, this is a topic we're we're facing at work right now of how do we how do we handle this new review burden and
You know, what I've what I've kind of been exploring is Bill, our engineering manager at Luma had put together a Laravel app and piece of this app like listens to pull request webhooks and it does an automated code review. The the first iteration was you know coming back with review and it like posts a comment on the pull request and then does a final recommendation of
changes requested or or or approved. and it the first iteration was like it it would get things factually wrong. you know, it just overall wasn't super helpful. And so this morning I just deployed a whole bunch of updates to it that I think will help make it easier. But I think what
The goal is, and what I'm trying to explore is let's try and get the AI code review to do a pretty good job of of overall review, touching on architecture, absolutely taking into consideration our rule files and what we encode in those rule files, and then kind of doing its path matching and everything to know when to apply the rules. but then
Pairing that with calling out specific areas in the code base that absolutely need to have human eyes on them. I'm talking like auth migrations, you know, maybe config changes, middleware changes, because you know, we we have stuff talked tucked behind.
Critical auth middleware, author, like, you know, authentication, authorization middleware. So we want to make sure that like route doesn't get added, you know, plainly and and not get tucked behind its correct middleware. so kind of just calling out specific places that need human eyes and then kind of leaning on AI review to to get that right. And then also kind of leaning on the
individual engineers process like my process before I open a PR I run my code review skill iterate on it then before I even get to opening the PR like so if we're doing that as a process like hopefully by the time the PR gets opened there's not a ton for the automated code review to find.
Part of the other piece of the automated code review is, you know, as a company we're using anthropic and cloud code. you know, right now the code review is still using anthropic, but I'm looking at changing that to open AI so that we get like codecs doing the like or a GPT model doing the code review. So there's an outside perspective, which mirrors what I'm doing in my personal life with with coding where
Anthropics doing the coding, Codex is doing the code review and then they're iterating together on that. I don't know, is this is this something that you feel like you're you're facing as well? I don't know.
Chris Gmyr (26:42)
Yeah, yeah, hundred percent. yeah, so the beginning of the year we basically got, you know, full access to all the AI agents, like club code, all all the things, and you know, the the amount of code going into our GitHub org has easily, you know, exploded. And it's put, you know, a heavy burden on people reviewing the code and also keeping even more context.
you know, in your head of like I gotta keep my own contacts and my own agents contacts and look at all of your contacts for, you know, the the team that you're on. And it's too much. And we're writing too much code to actually review like line by line as we did before. and because it's so cheap and easy for AI to do it, it's just sling in more code. And, you know, it's very easy to like glaze over, you know, any sort of, you know, PR at this point and just be like, yeah, you know.
Should be fine. I don't know. You know.
TJ (27:37)
Yeah.
Chris Gmyr (27:39)
And hope for the best, which is not a great position to be in. it might be fine for like personal, you know, side projects, whatever for a time, but like, especially in a a regulated space, like can't be can't be doing
TJ (27:49)
Mm-hmm.
Chris Gmyr (27:49)
that. So it's definitely something that I've started working on. so we had copilot enabled for the org, right off the bat. We've had that for, you know, a while now. And it's been okay.
But with Cloud Code, and then things working differently locally versus in you know CI and Copilot, they didn't really like merge together very well. It's like you would have rules and things like that locally that would work with Claude, and then you push something up and you know, Copilot wouldn't know anything about it. So in our plugins repo for engineering, I made a plugin for rules and
Converting those into GitHub instructions. So you can basically mimic your rules into copilot instructions. So it it's basically a hook that says anytime that the agent makes an an adjustment to your cloud rules, do this sync and push it over into this copilot instructions directory and basically mimic it the the content of it.
There's like a little bit different front matter and a little bit, you know, differences here and there, which is kind of annoying. But you do need those like physical files in there for Copilot to utilize that. So that I feel like helped. but it's still, you know, not really where I wanted it to be. So I ended up making my own we have a a shared GitHub workflow repo that people can conditionally like bring into their repos. And in there.
I was able to make like an AI bot reviewer using Claude and AWS Bedrock that will do the review for you and get a little bit closer to what you would get like locally with Claude. The problem with that is that we're still on like the GitHub flavored runners, even though like we're hosting those in our infrastructure as well. So a lot of times, depending on the prompt or the PR size or
what you would bring into the context. It would basically like blow out the contacts like very quickly. And it'd be like, too too large. Can't do anything with this.
TJ (30:02)
Mm-hmm.
Chris Gmyr (30:03)
Where like copilot, it would run, which is kind of annoying. So I've been like tinkering with that. It was working pretty well. I've updated it to Sonic version five. And I think it's on like medium effort. So I mean it's not
super thorough, but it does like a a decent enough job like finding gaps in issues. So sometimes it's definitely over zealous and just finding stuff to find things, I think. So I definitely need to tune the prompt that I'm sending it. but everything goes in as like one one round of processing. So I push through this like huge prompt for review.
I push through the individual diffs and do some like calculations along the way of like, hey, if we push all this PR through the contacts, like is it gonna blow it out? Or do we need to prioritize content that we are putting into the contacts? So we have like session logs, we have you know a bunch of tests, you know, things like that. So it'll degrade those like ratings as we fill up that context. So you're always gonna get the actual code.
But some of the maybe like documentation or the session logs or maybe some like large test files like might not go into there if you have like a really big PR. so it's something that I'm still tinkering with. I want to try like different models. I actually want to like offload this into like an AI agent reviewer inside of Bedrock that like runs differently and will be contained differently. and also be able to
give like a lot more feedback. So I want like comments on the individual like lines and files and you could, you know, plus one or minus one or thumbs up, thumbs down, you know, on that. And then I'll go into like the next review type of thing. But that has complications with that. cause right now the reviews are very like independent. They don't they don't crosstalk, they don't do anything like that. So like building that kind of stateful mechanism is gonna be tricky also.
but it's something that like I'm kind of driving that direction too. So doing a lot of tinkering and doing a lot of figuring out, but to the goal of we have so much code we cannot actually review all of this. We need to offload the mostly important bits to AI and the little bits to AI, and like you were saying, like I want to flag like
Yeah, look at the, you know, code, but this is the one spot that I want you to like humanize, spend five minutes and look at that. Like I want those flags like out in the open and being shown so that even if you do have ten PR reviews to do in the day, like you can only look at like this one part or you know, you have a little time and you know, review this whole other PR, you know, breaking that up a little bit. And also like, how do we get out of this like prose everything, you know?
So do we do like any sort of like visual design and a comment or you know showing like mermaid diagrams like automatically of like here's the you know previous code flow, here's the new code flow. so that like a human could easily say, like, that's actually not right. I don't need to look at the code. I am trusting the mermaid diagram for the flows of like, you know, past and current in this PR. So like
I don't know. What are some of these other mechanisms that we can kind of pull on and design that isn't just more pros and you know, terminal output and all that?
TJ (33:23)
Yeah, you know, and I I definitely lean on my PR message generation skill too, because that does kind of call out like here's architecture changes, is mermaid diagrams.
it also includes a section on like how to review this PR and kind of like steps you through what files to look at and like in what order to like make sense of of all the changes. So that's kind of helpful for a human, but yeah, I think I think it's I I think it's trying, I think the goal is to get to a process where we can lean on the an AI code review.
for the majority of a PR and then call out the four human eyes specifically pieces that are like critical pieces that need to have manual review on them. I think that's I think that's the goal of where we're heading to. I'm not super happy with the current solution. It's something but it it needs to get better. you know I think
I think there's one one solution I might want to look a little closer at. I don't know if we would use, but I wanna investigate a little closer. Code Rabbit. I believe I believe it's Code Rabbit. they were at Laricon as a sponsor. I've I've seen them around for the last couple years.
You know, and that's that's like an AI review solution. so I think we'd we'd I'd maybe look to something as that to like maybe learn some lessons or get some ideas from. but yeah, I think I think the goal is to get confident enough that we can lean on the AI review, call out specific spots that need to have human eyes on them. And I think I think that's the way forward.
Chris Gmyr (35:06)
Yeah, 100%. I've looked at Coderapit before for personal stuff. I don't have to look at it again. Cause I do the adversarial reviews too and also trying to build my own set of review skills for my personal projects. But it is so hard to find the balance of like, yes, you found something, but is this actually important? yeah,
TJ (35:31)
Right.
Chris Gmyr (35:32)
we've we've talked about that before.
Like just because you found something doesn't mean it's an issue or it's an issue right now. It's like, this side project, like I'm just using it for me and I'm gonna use it the proper way, you know. I don't care about these like ten other random edge cases that something could happen because it's not going out to a thousand or a million people, you know, where people are gonna find that. So it's like adversarial by nature is adversarial. Like it's it's built to find something no matter what.
And like it's very rare that it's not going to find something unless the code change is very small. So then what happens is that like it finds something, it's like, yeah, that's a valid you know, call out. Let me fix that. And then the next thing, it bypasses the thing that you just fixed, but now it finds issue number two. yeah, issue two makes sense. We should fix that. And then now this like little itty-bitty change has exploded into
five, ten different commits and it's still going and still being adversarial and finding these like little nitpicky things. And your scope has blown up 10X. Your PR has blown up 10X. And now it's just like my one, you know, little change to a component or even you know a design or whatever the case is. It doesn't even matter. It's like it's blown completely out of proportion and handles all these like 20 different edge cases. And it's like, why? Why can't we just
ship it and then if an issue comes up, then we'll fix it then. And you know, some things like it does find and it's very beneficial, but like where is that line? And again, how to like tune tune the reviewers, but also like the the workers who have to review the reviewers work to say like yes, I should implement this fix because the reviewer reviewer said so.
TJ (37:22)
Yeah, I that I was literally just wrestling with that yesterday of like it just kept on finding things no matter what. And it's like, how do I how do I prompt its way into, you know, where that line even is? And like how do you define that? And geez, like it's it's it feels really nuanced and cause
Sometimes like sometimes I do care about a guard clause. Sometimes it really doesn't matter and you're just trying to find something to find something, you know? Yeah. Yeah. It's it's a tricky one. but that's it's definitely like it's it's a big topic at work right now and
it's gonna continue to be because yeah, it's it's unmanageable now. Like we've got people spending whole days, their their entire work days doing code reviews. Like you've burned an entire day of doing anything other than code review. And that's not acceptable. Like we can't be doing
Chris Gmyr (38:19)
Yeah.
TJ (38:19)
that.
Chris Gmyr (38:20)
Yep. Exactly. And the other thing that's been on my mind is like how to test and tune these things. So with like my AI review script bot, whatever, I do a lot of like manual testing and I'll run like the the dev branch on like repos that I own so I can be like, you know, how does this feel compared to the old way of doing it? But what I've
realize now is like I need to make like a whole testing framework and comparison and like eval loop for like all these things. It's like take, you know, basis or commits or PRs, but like bring them in as like demos and diffs. And, you know, now I need to run this across like, you know, Sana at different efforts and Opus at different efforts and see what actually runs in the GitHub runner.
and see what the output is and see if it's too fiddly or too aggressive or not aggressive enough and like all that like human in the loop stuff to like tune the the agent, the prompts, the mechanisms to do all the things, like that's also a whole ton of work for a team and company who is not responsible for code reviews, you know?
TJ (39:31)
Yeah,
I I Claude and I built an eval harness as part of the work because that's that's just like every AI feature that I have built at Luma for the platform. I build eval harnesses to go along with them so so that Claude can iterate on its own and and kind of like run the eval, tweak some things, run the eval, tweak some things. and so we we did a couple of those iterations, but like I also
need to try to find confidence in the evail harness too, right? So it's
Chris Gmyr (40:03)
Mm-hmm.
TJ (40:03)
it's also it's also balancing like how much of my time do I spend on like automating this versus getting other work done too. So yeah it's it's tricky, but I think we'll we'll slowly get there as we iterate. But yeah, it I think this is a a big problem for a lot of orgs right now.
Chris Gmyr (40:21)
Yeah,
yeah. A hundred percent. So yeah.
TJ (40:23)
Yeah.
Well, on that note, man, I think we're probably good to wrap up. I think we both have some stuff going on here at the next few minutes. So
Chris Gmyr (40:29)
Yeah, let's wrap up.
TJ (40:31)
cool, man. So
Let's see. Thank you all so much for listening to the Slightly Caffeinated podcast. Show notes, including all the links of things we mentioned in social channels, are down below and are also available at slightlycaffeinated.fm. If you have any questions for us or content suggestions, go to the ask a question page on our site and we'll feature you on an upcoming episode. Thank you so much for listening. We will catch you next week.
Creators and Guests
