Time to get back to Assembly, I guess.
But do you trust the CPU?
Well, fuck
The CPU runs Minix, are you going to tell me you don’t trust Minix now?
deleted by creator
I think about this every day.
Dev’s what?
Dev is screwed.
Maybe? If you poison the prompt then there’s evidence and it can be undone. Poison fragments of the source training data, however, and that’s some KT shit right there. Enterprise foundation models cost bonkers money to train and pretty much slurp up all the data on the internet for mostly automated annotation. Stick something in an obscure part of the internet which becomes part of the training and produces the malicious response and it’s going to be both hard and expensive to detect or correct.
Should be relatively easy, with the amount of once trusted packages that become attack vectors

except for poison to take in, it should be a pretty significant part of the dataset. Also, ngl, i’m not much informed on the topic, but aren’t all the datasets, if we’re talking about generic diffusion models and LLMs, already been formed? From what i gather, the innovation in AI mainly comes from utilizing new architectures, rather than training a model on something unique.
The datasets are constantly expanding as new content is generated online. There’s a degradation issue currently where the models are training on incorrect data generated by previous iteration of their own or other models and effectively poisoning itself to more confidently give the same incorrect information in future.
i’ve heard of the dataset poisoning and degradation caused by llm-generated content present in the dataset myself, but i’m not sure whether it was a practical observation, or a mere experiment. And I still fail to see how new datasets are really useful for developing a new llms, or how it’s a problem for the devs to switch back to the older datasets.
And the cornerstone stays the same: to have any significant effect on the final LLM quality, shouldn’t the poisoned (either by llm-produced content, or by intentional poisoning) data portion be… well, statistically significant?
One thing to keep in mind is that, when it comes to LLMs, the models have not significantly changed in architecture.
There’s been new experiments and advancements in architecture on neural networks, and machine learning for specific applications. But LLM, as they are being commercialized by AI corporations to the general public, have stayed relatively the same. Except for one thing. Increasing in size. Larger datasets, or more specialized datasets like with coding, and larger number of tokens in memory. This is why it takes such large data centers. It’s all been just brute forcing greater capabilities by enlarging the models.
One of the things with LLM is that all the dataset influences the weighs and probabilities of the results. Even if the dataset includes a single event of a chain of words (think of the pizza with superglue incident), it can show up in the results eventually.
source?
https://www.nature.com/articles/s41586-024-07566-y
LLMs already tend to be the average predicted output for a given input; training on LLM data makes this worse and causes them to become less varied, less dynamic, more towards the mean generated by previous models, and more likely to spit out hallucinations
For the uninitiated: who compiles the compilers?
Quis custodiet ipsos custodes?
I’ve had it on my todo for years to work through ddc and trusting trust.
Which is a method to verify a compiler is matching its source and thus trustworthy.An orthogonal approach is reproducible builds, which among many benefits can make sure a few people verifying things benefit everyone who can then see they have the same verified binaries.
Something you might like: https://bootstrappable.org/
The very fact that Anthropic is now injecting a kind of watermark into every output, is solid proof that such a Ken Thompson hack is a inevetable risk
The ending is rather unsatisfactory.
spoiler
Some how like all sophisticated technical stories about ai it ends with the acceptance that there is nothing we can do because the alternative would be going back to analog.
Sounds like quitter talk to me. It’s not like AI is an evolutionary process that just happens somewhere, it’s a localized tool - and even if it spreads itself, that would be by far the most humonguous virus ever. I think those authors just write down their own fears.
lol wow this is gold
Who??
What does that even mean. Whoever said that just uttered some empty but smart sounding catch phrase. Such is all the talk about the wonders of Ai
Linus Torvalds
Not exactly, but something along those lines.
We are talking about a guy still using mailing lists and patch files to conduct development on one of the largest codebases in the world, not exactly someone who jumps on any new shiny thing just to sound smart.
https://thenewstack.io/torvalds-ai-programming-productivity/
just to clarify, he’d not so much called LLMs “the new compilers” as he compared both to each other in a sense that an LLM is just another layer of analysis tooling between the developer and the final machine code, which, IMO, sounds much more reasonable than calling LLMs “the new compiler”.
I think it’s supposed to be that how AI turns high level instructions into code is compared to how compilers turn code into assembly. Implying that using AI is just a natural extension of the handing off work to the computers that we’ve already been doing.
I wonder if you gave different AI models some c++ or something and told them to write assembly based on it how well they would do compared to an actual compiler
They already have a solution to the Trusting trust attack thanks to DDC and live-bootstrap’s work.
https://guix.gnu.org/en/blog/2023/the-full-source-bootstrap-building-from-source-all-the-way-down/
Mmh… but this concept only fully works on Open-Source hardware, doesn’t it? Otherwise the microcode or CPU itself could still be an attack vector to infect the bootstrapped compiler?
(Genuinely curious, I have no clue)
That would be the deeply classified Nexus Intruder Program that no one has yet solved. Which would be hardware that infects software that in turn infects the next generation of hardware.
It’s a tool, use it where it works and don’t where it doesn’t.
That saying (or, dare I say, thought-terminating cliché) glosses over the consideration of risks and costs that result from its uncritical usage. No sane person should trust it without serious reservations just because it fits the purpose, and that’s true whether or not it bears the latest combination of letters peddled by tech bros. It’s a dangerous, irresponsible mentality. It’s a tool the same way a sledgehammer made out of plutonium is a tool.
Plutonium sledgehammer sounds fucking awesome though
Pretty effective because of its high density, too. It’s a shame it’s brittle and would give you heavy metal poisoning.
As ways to die go that also sounds pretty rad
Oh, that would be rad alright
Unless you pick a particularly short-lived isotope, the radiation wouldn’t have time to hurt you before you died of regular heavy metal poisoning like you’d get from something boring like lead or mercury.
🤘
I mean in an old nuclear bomb that’s pretty much the situation. Use a conventional explosive to slam some plutonium like a sledgehammer and nukes have proven to be a strong deterrent when a country had them. Still a tool with a use case
Your comment is a sweeping statement that is thought terminating itself.
No sane person should trust it without serious reservations just because it fits the purpose
Using a tool for something it’s intended for automatically involves some thought process. And this thought process also involves its limitations and potential for danger.
Your criticism stems from your dislike for AI and your opinion that it’s generally not useful. An opinion that is shared widely here, but with little to no proof at all.
And then your boss says: “Use it for everything or you’re fired. Why are you using less tokens than anyone else? You must be more productive or else…”
People probably say that about asbestos too.
And they’d be right
Of course, it still has uses. Just not in home insulation.
Your comment is so fit for purpose, displaying such a massive willful ignorance about the world in general. Because the biggest problem with asbestos wasn’t using it for home insulation. It killed millions of workers in manufacturing plants and mines. Its biggest industrial use case was for electrical insulation of high voltage transmission lines. Where it also killed hundreds of thousands of technicians in the us alone. Before that it was used for lamp wicks. Where it killed at least a couple of millions more over a century. Today, asbestos kills 250 thousand people a year, 12 to 15 thousand in the us every year.
I supposed any tool can be considered useful if you’re ignorant enough about it or if you’re willing to lie without remorse.
It’s a tool, use it where it works and don’t where it doesn’t.
But that doesn’t work with the people creating AI as they are totally dependent on the believe that AI can do absolutely everything (and an artificial general intelligence is just moments away…) to justify they insane investments. So they will make up a million stupid narratives why some people are “actually” not using AI as it obviously can’t be because of AI shortcomings…
If the tool don’t work then don’t use it. I wouldn’t try to unscrew something with chopsticks
You are probably also one of those insane ideologues that refuse to hammer in a screw for some reason, although you know how well that hammer worked on nails… 😂
AI or not to AI is the same stories I heard from old time machinists when cnc productivity came for that industry.
not really related to the post, but NGL, it kinda fascinates me, how with the advent of LLMs, they became the ultimate punching bag for whenever something works bad, as if people didn’t write even more horrible things without them (my regards to javascript and python).
Will share the down votes with you, because I think the same.
Claude opus is better than a significant amount of people I’ve worked with, and it’s still significantly cheaper when now. Mind, we are not professional developers, but we need to develop a lot a DS solutions.
Even if AI was not hot garbage the other [insert number] % of the time, would that equate to better material outcomes for actual humans? Not under capitalism, that’s for sure.
What even is a “professional developer”? I mean, given that you aren’t an egg-headed RnD developing new technology for someone like Nvidia, the job of a developer is 95% mind-numbing routine and typical tasks solved by applying ready-made patterns.













