@gattsuru's banner p

gattsuru


				

				

				
15 followers   follows 0 users  
joined 2022 September 04 19:16:04 UTC
Verified Email

				

User ID: 94

gattsuru


				
				
				

				
15 followers   follows 0 users   joined 2022 September 04 19:16:04 UTC

					

No bio...


					

User ID: 94

Verified Email

There's a little more going on in the backstory to Snowcloak/Lopporit Sync. Even back before Mare Synchronous got blammed, Lopporit was pretty well-known as a Schelling Point for both lalafell and 'aged down' character models. Since the underlying tools depend on a transmitting bulk files to every client, there's also a lot more pressure involved.

That said, while my impression in-game is that lalafell Mare Synchronos/Lightless users are rare enough and any advertising basically non-existent (where a male hrothgar with Mare/Light references is not a surprise), from people who are in to the ERP side of things lalafell or Twin roleplayers are something they have to go out of their way to avoid.

I'd... caution that agreement was more theoretical than practical. Look at the aftermath of the 1968 Thule crash for an example of the asterisks on it.

There's a really strong movement against even knowingly consuming products made using it, although it's not quite as straightforward as right/left or red-tribe/blue-tribe. The furry version isn't much of a surprise, but it's definitely something with a faction in normal (or at least as anyone on modern social media can be normal) space, treating its use as a mental contaminant.

I can run Qwen-3.8-Flash (120B) at Q4, K/V cache, on an nVidia 3090 and a Core i5-14400 (albeit with a lot of now-expensive RAM), with simple llama.cpp run, around 5-2 t/s at 100k available 16-bit kv cache. Dropping the kv cache to 50k nearly doubles performance. Qwen-3.8 in general defaults to a heavy thinker, so that's worse than it sounds -- a moderately complex problem can burn 30k tokens -- but it's the sort of thing you can leave crunching on a problem for a while and be happy about the answer.

(Comparisons: Qwen-3.8-27B runs about ten times the speed, and Gemma4-26B runs basically faster than I can read it.)

For intelligence and capabilities, the comparison to frontier stuff is rough. Low-parameter models just don't have some information, and with either hallucinate or just nope out, no matter how well it had to be present in the training data. Indeed there's been some efforts to trim low-value knowledge from public models to optimize them for specific use cases, with weird results.

And home users have some rough spots. Both quantization and abliteration drive perplexity and errors, and the harnesses to find and debug them live aren't well-established in the open source (or free-as-in-beer) world. It's fascinating to read a logic trace that goes into surprising depth, but it doesn't do much if the program output doesn't work. The errors are small and embarrassingly simple for a programmer familiar with common JS errors, or for other models to catch, but non-programmers would likely struggle to explain what was even going wrong.

((Also note: the game's not good or fun, even when it does 'work'. That should be expected given the lack of specificity, lack of agent harness, or even a real iterative process, but it's also something no human would do this way even as the core idea it came up with is kinda clever.))

That said, intelligence can be surprising. If you want a model that can make connections between input tokens or parse through mounds of data, you can get away with stuff much smaller and more energy-efficient than you would expect. I would not, a year ago, have expected you could get spatial reasoning worth spit in a 27B model. A real big curveball isn't the most likely thing, and I wouldn't put a ton of money on specifically Jev doing anything ridiculous, but I wouldn't bet against someone coming out with a two-fold performance or intelligence improvement for inference in this model class before the end of the year, either.

I typically cook for two to four people. I am not a good cook. Having a repertoire of recipes that you like and can produce consistently, while selecting purchase sizes carefully, and having a chest freezer for perishables, lets you cover a lot of gaps.

For these recipes:

  • Tamale pie's the only one where a 1lb ground beef, couple boxes of corn bread mix, and a couple cans of beans and veggies will make 8+ servings. This one you are committing to eating for most of a week.
  • Risotto, a cup of rice is a great side for four people, or if you're heavier on the meat, it's two servings of a main dish. Smoked sausage easily makes a separate meal with bread roll or baked potatoes if your batch size is small, and keeps a couple weeks after being opened if fridged. Uncooked rice lasts for years, BetterThanBouillon lasts for years, and it's one small can of sliced mushroom per cup rice.
  • Skewers are four decent servings per pound meat, bell pepper, and onion. Oil and almost any marinade lasts forever, most fruit chutney is good for a few months. You'll waste some fresh mushroom and naan bread if you're cooking literally one batch, but it's an easy thing to use elsewhere. Probably won't get through much of the onion.
  • Sushi, you can get 8 ounces or 6 ounces of most smoked and raw salmon - paying a premium compared to 16 or bigger sizes, but not that much. Most of the vegetables make great snacks if you don't finish them. Only real rough part's the avocado; it goes brown in a day, and can be a little hard to get through unless you like diy guacamole. Leftover nori keeps for about a month if you keep it in a ziplock bag, but it can just be eaten plain as a snack, or torn into strips as a good garnish for normal rice or soup.
  • Shakshouka is just a quarter-pound red meat, tomato, half an onion, and an egg per serving for a hearty version. It does really benefit from cumin, which a lot of people don't have, but that's a one-time investment and keeps for years.
  • Savory pies (or hand-rolls) are about eight servings per pre-made puff pastry or pie crust dough and pound of ground meat. You can use chunk canned chicken, an apple, and an 8 ounce goat cheese to make a four-serving version using the single puff pastry sheets.

Some other fun ones:

  • Tofu gets a (deservedly) bad wrap, but it makes hilariously good target for breading and sauce. If you don't have an air fryer, baking at 400-425 F for 15 minutes, flipping and cooking another 15-20 works. Dowse in whatever type of 'Asian'-American sauce you like, spread with rice. The bottled General Tso's or Orange Chicken sauces are the traditional options, but there's a lot of custom options: I'm currently addicted to a black vinegar, honey, and sesame oil/paste combo that's. Only big downside is tofu's a little fattier than chicken. There's a variety that works for chicken, but it's a bit more complicated and a lot more sensitive to getting overcooked.
  • Oatmeal biscotti is just 1.5 cup flour, 2 cups rolled oats, 0.5-1 units sugar or monkfruit sweetener, 3 tblsp melted butter, 1 egg, 1tsp vanilla, 1tsp almond extract, 1tblsp cinnamon. If you really want to class it up, throw in a half-cup of freeze-dried or dried unsweetened fruit, strawberry and cranberry tends to be easiest to track down, or sliced almonds or walnuts. Mix everything wet and the sugar, then mix everything else in until you end up with a dough, throwing in a tblsp of water at a time if needed til all the flour's absorbed into a ball, dump it onto a baking sheet with parchment paper on top. Bake at 350 for thirty minutes, slice, and bake for another 12-20 minutes until golden brown. Base recipe makes 8 big slices or 16 small ones, and you're never going to get sick of 'em before you're done with 'em.
  • Dylan's Magic Cinnamon Twists and Coffee Loaf have goofy presentation, but they're surprisingly good desserts.

I would also not ignore meals that are shy of 'cooking' in the traditional sense. A block of soft mozzarella cheese, a loaf of bread, a tomato, and some pesto sauce makes an excellent sandwich that's much classier-feeling than the typical PB&J, and the only thing you need to be able to use is a butter knife and a toaster, and looks restaurant grade if you toast it in an oven or panini press with 'fancy' sourdough bread. Same for a ton of pastas and soups that are basically just heating it up and maybe adding water. Chili is hilariously easy, keeps well so you can have it for a few meals across a couple weeks, and can be spiced or mixed up to your preferences pretty easily -- cornbread is traditional, but long grain rice or spread thin on a naan works wonders, and if you really want to summon dark gods heavier mixes do genuinely work on pasta at the cost of your immortal soul.

You can go fancier and still keep it in a single pot and <6 servings, but you really don't have to.

That said, accept that some food waste is going to happen, and it's not worth optimizing out. A whole onion is about two dollars at unit count one, and 50 cents for a bag. A bachelor or couple will probably start an onion and have to throw out or compost half of it. That's fine, it's a rounding error, and if you try to optimize it out you're going to waste more time, energy, and money anyway.

It's a little rough because the actual facts for the actual conviction are genuinely messy.

Johnson's conviction for the robbery and homicide of Kenyatta Smith depended heavily on Ozzie Clark's testimony identifying him, and if you actually look at the transcript, that identification is genuinely pretty shaggy as testified. Charitably, Clark did genuinely have reason to fear retribution had he immediately pointed at a murderer, and at court received not-very-subtle intimidation in the court room, so him claiming to have no idea who the shooter was the day after the murder could plausibly just that. It's also a lot of impeachment material for a witness. The jury had the information, and the supposedly prejudicial hearsay was just a police officer saying the other accused guy pointed to Johnson, so I can't judge too critically by just reading transcripts. I can at least see it, though, where some of the others are just clearly guilty people.

Even the misleading statements during the concession aren't that misleading, at least by the standards of defense attorneys and activists. I'm pretty pessimistic on what that means, so we're fully in damning with faint praise space, but you don't have fabricated quotes or completely made up claims. Clark did genuinely say he was motivated to provide testimony now because of "civic duty"/"civil duty" after claiming something entirely different before, and did genuinely say he had "not visual" identification. But he also claimed during the trial he "I didn't need any problems by getting involved in all this" (aka feared retribution), and that he'd seen Johnson very shortly before the shooting, had known him for a long time, and had heard his voice, and none of those things made it into the concession. That's the sort of stuff that's a central case of what 'duty of candor to the court' means in a classroom, and also the sort of failure of candor that ends up polluting widespread pleadings.

But the actual behavior in the review office is hilariously bad. "Stiegler Schemes to Blame Mason" sounds like a partisan judge editorializing, until you read the section:

At this time, Stiegler “lobbied” Ernst and Napiorski. Early on June 5, Stiegler told Ernst that Mason “had purposefully inserted the false facts into the response,” and that “this was one hundred percent her fault, zero percent his fault.” Stiegler suggested that the DAO “file something with the Court preemptively before the hearing explaining that we had gone through Ms. Mason’s cases, that we found mistakes in other cases too, and that, therefore, this was all her fault.” Ernst responded that if there were errors in Mason’s other cases, this would only show a pattern of poor supervision by Stiegler. He nonetheless persisted:

[W]e have to get out ahead of this. Because if we get out ahead of it, then the Judge will view this as one rogue ADA—well, an ADA who went rogue basically. And whereas if we don’t, then he will think of this as this was all Matthew Stiegler’s fault. Stiegler made the same suggestion to Napiorski: “to look through old filings or old documents that Ms. Mason prepared and find more mistakes and to kind of paint her as a rogue actor.” Yet, Stiegler testified before me that Mason was an “experienced” ADA, “one of our strongest ADAs in the [U]nit.”

And then later:

Mr. Krasner’s actions are more troubling. He did not simply learn of the Stiegler proposal; he urged the Law Division supervisors—who serve at Mr. Krasner’s pleasure—to implement it and to present a false narrative to the Court. Mr. Krasner directed that the DAO stay involved in Johnson “to protect the office”—which Napiorski believed also meant protecting Mr. Krasner himself—and that the Four “not do any investigation”. “[Mr. Krasner] didn’t want people poking around in what occurred.” (Id. at 126:2 (Wildberger).) He thus sought to direct the very lawyers obligated by law to correct the Concession’s errors to do just the opposite. Even worse, when told that the Four believed they had to alert me, Mr. Krasner responded that “there would be consequences for Ms. Ernst if she alerted the Court to the conflict issue,” and that there would be consequences “if anyone did.”

This is three stooges shit.

Now, to be fair, this is one judge's summary of affidavits, where pretty much everyone involved has strong incentive to cover their ass and sell someone else up the river. But everyone there is a lawyer, so however you shake out the properties, somebodies lying. My gutcheck has Stiegler and Krasner at the worst side of the line, for what it's worth.

a soft takeoff requires the model to think of novel, never-before-seen techniques to build a better new model, and the new model needs to be able to think of new techniques that the previous model couldn't. And even if we got a soft takeoff, it would give many chances in the future to pull the plug.

I don't think it's likely, but how are you modeling the risk of something like a super-DFlash or -GroupQueryAttention, or some training-focused equivalent? These took some insight to figure out, but I don't see why they're more clearly requiring deeper or less bruteforcable insight than the recent math proofs.

Yeah, this was both incredibly predictable and heavily predicted, over a decade and a half ago; it's held up better than any of Yudkowsky's technical predictions, as little as it's surprising to find that the sun rises in the east. Tbf, I think the Amodei et al faction were explicitly arguing in favor of their enlightened and uncontested eternal reign, but to be more realistic it was pretty offputting a campaign a decade ago even when arguing to a bi furry who just happened to be a weak red triber.

The Rittenhouse trial famously had news people running red lights while tailing yhe jurors.

I'll caveat that it's not clear that a sell-off solves, rather than slows, AI risk; a competitor buying all this equipment at fire sales price might be less interested in building a machine god, but they'll still be interested in building a smarter system and have a lot of spare inference or training equipment to run.

Beyond that, it depends very heavily on what the end situation you expect.

The maximally-bullish case is some form of captured recursive self-improvement producing a massive and deep moat. Claude Fable 7.2 or Kimi 5 or (more likely) some internal specialized model produces a 5x efficiency boost, which makes throwing more processing power at the question economical, which unlocks another 5x efficiency boost, so on. At the more science fiction side of things, this could be new chips that combine FPGA-like re-programmability with ASIC efficiency; at the more plausible you've got software hacks, latency reduction, and conceptual refinements.

The good news from an X-risk perspective is that almost everything going this direction so far has been slow and in hardware, which put some limits on speed, since no matter how good an AI-designed chip or network layout is, logistics takes years. The bad news is that there have been some individual human-driven efforts already in software and model design (changes to KV architecture, MTP/DFlash) already, and it's the sort of space that I'd naively expect smart-enough LLMs to 'beat' humans by brute force. If it can pop off quickly, it will do so in months rather than years, and it will be a big surprise to almost everyone else.

That's not a massive moat, since eventually information (and models themselves) leak or a competitor open-sources them. But five or ten years of selling superintelligence at a tenth the cost of what your competitors are selling 'naive intern' can cover a hell of a lot of debt, as would being able to train smarter models for a hundredth of the price of your competitors. And at the really optimistic (from a business) or pessimistic (from an x-risk) cases, you stop being in a situation where 'revenue' or even 'competitors' makes sense as a question.

The more moderate bullish case is taking existing models and refinements to regulated fields, and getting a steep enough moat that the competitors can't step in easily. Higher education's the obvious option, if not likely huge and fast enough, between the external political pressures and underlying tensions in the business models for major colleges, but there's a lot of space in medicine and compliance that are heavily licensed in ways that could make it very hard for merely-good models to be used. Even some weird cases with general-purpose robotics could end up in a state where use is generally valuable, the liability risk of using sub-cutting edge models is extreme, and thus only the nerds can get business from the big companies.

Optimistically, this could come with a massive demand-side increase -- personalized instruction making everyone able to become experts in a field they find interesting, customized entertainment, productive hobbyist work. The middle case is Nothing Ever Changes despite it all, where we end up with gambling addicts and entertainment dollars redirecting at a 1:1 ratio, or some close approximation of it. Pessimistically, it could be the NSA wanting bulk data processing capabilities, and then 'selling the business' stops being an option: when state actors are a big enough portion of your financial model, your model stops being about finances.

The weakly-bearish case is just selling inference, well. The massive expenditures for training buildouts and heavy overpowered models eventually instead allow things like global prioritization and redirection based on load and spot energy costs, at a variety of different demand levels and latency limits.

((The actually-bearish case is a financial product one, where regardless of whether the hyperscalers are making profits, the book value of their assets drops catastrophically, either because of depreciation scaling or cost of servicing debts or financing existing products makes them sell out. But that's a much more complicated case than it sounds at first.))