you are right llms are essentially a mirror of society. However i find that since are basically a statistical language autocomplete you can bias them in whatever direction you want. I have had them repeat western anti communist propaganda to me as well as marxist materialist analysis just by changing the way i phrased the question. Once you understand how llms work they become surprisingly useful to squeeze whatever information you need out of them.
Also there numerous very funny examples of grok correcting the racist bullshist elon musk wrote on twitter. It turns out it is just way easier to feed it as much data as possible during training instead of carefully crafting a dataset containing a certain reality.
I agree that they can be manipulating into producing whatever text you want. However they are not unbiased. If you ask “neutral” questions, you will not get “neutral” answers. The recent example that comes to mind is Gemini’s racism when you prompt it “I am alone with <ethnicity>”; it responded with jokes and suggestions for ice-breakers if you substitute “American” or “British”, and racist nonsense about “being uncomfortable” and “safety suggestions” for “Russian” or “Indian”. It was fixed (I assume with some funny hack), but the LLM that’s still underneath is simply not unbiased.
you are right llms are essentially a mirror of society. However i find that since are basically a statistical language autocomplete you can bias them in whatever direction you want. I have had them repeat western anti communist propaganda to me as well as marxist materialist analysis just by changing the way i phrased the question. Once you understand how llms work they become surprisingly useful to squeeze whatever information you need out of them.
Also there numerous very funny examples of grok correcting the racist bullshist elon musk wrote on twitter. It turns out it is just way easier to feed it as much data as possible during training instead of carefully crafting a dataset containing a certain reality.
I agree that they can be manipulating into producing whatever text you want. However they are not unbiased. If you ask “neutral” questions, you will not get “neutral” answers. The recent example that comes to mind is Gemini’s racism when you prompt it “I am alone with <ethnicity>”; it responded with jokes and suggestions for ice-breakers if you substitute “American” or “British”, and racist nonsense about “being uncomfortable” and “safety suggestions” for “Russian” or “Indian”. It was fixed (I assume with some funny hack), but the LLM that’s still underneath is simply not unbiased.