• 6 Posts
  • 556 Comments
Joined 3 years ago
cake
Cake day: August 29th, 2023

help-circle





  • (2) has a big “if” in there

    Its a super big if, but it is one lesswrong (and Anthropic, with their “model welfare” pseudoscience) is allegedly seriously considering as a possibility (since, you know, they think LLMs are already AGI and Eliezer was accusing AI-dungeon, you know, GPT-2, of deliberately scheming). If anything, it speaks to some possible hypocrisy/motivated reasoning that they consider LLMs AGI but not possibly morally relevant.

    conceptualized capitalism

    This seems to be a lesson most lesswrongers have had rubbed in their faces repeatedly over the past 5 years with everything about LLM company’s behavior, but they just don’t want to learn.







  • I believe the main ingredient is Lean, which is a formal language resembling a programming language. Math proofs written in Lean can be verified deterministically with a computer, which really helps mitigate the hallucination problems of LLMs.

    100% this. Also, looking back at an earlier example that was actually written up in more detail, AlphaGeometry 1 got 28/30 problems, but entirely stripping out the LLM from the system, the symbolic logic proportion alone could get 14/30, and replacing the LLM with different heuristic methods could get 18/30 and 21/30 (for different methods).

    Even if math research works out perfectly well (which is a still big if), it’s not going to pay the bills. They would need to find a use case in the real world, where hallucinations can cause serious damage and cannot be formally prevented. And they have certainly tried. Math will not change the fact that all of this will collapse.

    The boosters and LLM companies still believe LLMs get their current level of performance by generalizing and not just memorizing facts (and maybe a wide shallow pool of weak heuristics). So they are hoping by pushing the LLM performance up in some narrow domain they can churn out synthetic data for, they will see some large general improvements in LLM performance.




  • I had two main guesses even before reading into the details…

    The obvious, they set this up for the doom crit-hype value. That is how basically every one of these stories like this to come out of Anthropic turn out to be once you read the details of the setup. The lack of details is to keep the mystique and hype up.

    The other guess… their coding agent actually did manage to vibe code its way out of its ‘sandbox’ and past Huggingface’s security, but the security on both ends (OpenAI’s"sandbox" and huggingface’s code) was itself laughably bad vibe coded garbage. The lack of details is to obfuscate just how garbage they are. I think this guess is pretty well supported by the linked Mastadon posts (although I still wouldn’t discount my first guess until more details are pried out). My “favorite” quote:

    “prepend some text with processing instructions and an SGML tag, and then hand that blob to an llm.”

    I fucking hate this timeline.

    and they turned it into a mutual marketing stunt

    Their marketing actually seem to be trying to take different spins on it…

    HuggingFace stressed that it worked out the attack was going on using an open weight model it hosted itself — because the guard rails on the commercial model HuggingFace tried wouldn’t let them do security work. So you should use open weight models from HuggingFace.

    Yeah, its an opposite spin than Anthropic and OpenAI have been trying lately, where they want to shut down all open weights model because something something China something something too dangerous something something regulate our competition out of existence. (It is funny, some of the boosters on hackernews and /r/singularity are actually acknowledging without an artificial moat to help boost OpenAI’s and Anthropic’s capex for bigger and bigger models will stop. Even though they still deny Ed Zitron’s calculations of the financials.)

    Anyway, I guess even if they have some very opposed strategic directions, huggingface and OpenAI are aligned on maximizing the hype.


  • You know, given automated proof checkers, I was naively assuming mathematics was one field that gen-AI would have a hard time screwing up. Even programming is too difficult to write thorough testing for. But a proof (or counter example to a conjecture) seems like it would have to be solid if it passes lean or whatever system for validating it.

    But the threats #3 and #5 the declaration lists make me consider the bigger picture. Academic fields without clear capitalist payouts are already underfunded and under respected. Pure mathematics could, at least up until now, draw on the respect STEM gets, but with math proofs getting used as fuel for the LLM hype machine, there are a variety of unpleasant ways things could twist.

    Threat #1 makes me wonder… if LLMs+formal verification systems manage to pluck lots of low hanging fruit, and we are left with harder stuff that not enough literature exists as training data for LLMs, it seems like the entire educational pipeline for producing mathematicians could end up screwed up.




  • That was kind of a nothing post? I think the author is misrepresenting the growth (or at least failing to connect the most important dots) of the lesswrong rationalist by portraying them as a natural evolution of the skeptics movement and not the deliberate cultivation of an audience by Eliezer looking for people to spread his ideology to and by Thiel looking to influence farm and create a cult incubator. And then the post is titled “my falling out with the rationalist”, and they write they "wanted to offer a more personal perspective " but they don’t actually discuss that much of the personal angle either…