RU version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
83% Positive
Analyzed from 1444 words in the discussion.
Trending Topics
#gemini#claude#better#more#model#gpt#code#google#students#writing

Discussion (53 Comments)Read Original on HackerNews
The length of the response is a huge factor:
> Students tended to prefer longer responses. The selected answer was 37% longer on average than the alternatives. The longest response won 47.7% of decisive writing comparisons. The shortest still won 25.0%.
So the score is partially a proxy for longest responses.
Makes me wonder how much the reviewers actually read the text. Were lazy evaluators picking the text that looked the longest or most structured without reading it all?
Interestingly we trained a simple LORA layer on top of Inkling and it turns out a model can smell which model wrote a response 70% of the time. I wonder if the smell of gemini is just more preferred by students.
Given that this is increasingly the go-to for a college degree, college needs to rethink its cirricula and place in the world. Or at least get rid of the essay.
Lest it become a place where student and teacher ais go to play pay-to-win social deduction video games.
They're adding steps like having the students discuss and defend their essay, which immediately reveals the people who had AI write something and thought they could bluff. This triggers complaints about social anxiety and such, which are unfortunately becoming the go-to defense when unable to discuss the work.
They're also moving toward more in-person writing. Instead of long essays, shorter writing segments as part of the test. Submitting a written essay earns you feedback from the professor and a better understanding of the topic, but that's it.
When prompted to use simple English without jargon, it's still filled with load bearing honest caveats in every footgun seam it talks about — what I should have led with. <Insert whatever other Claude cliche you prefer>
And I'm not the only one to notice this. Next time it's up for renewal, my team is abandoning it for GH Copilot in order to use literally any other frontier model.
I do go back to ChatGPT for things that need to be calculated, but Gemini simply hallucinates vastly less.
I did get suckered into extending my Grok subscription which is really fun for images/videos. Grok in car is a must too (my jaw dropped when I asked what is pantsula music and it added multiple albums into playlist for me).
Given all that, I can definitely see how Gemini would be preferred. And good for Google, because I would rather my offering be the top choice for the most people rather than better serving a small subset of users.
One could argue that Google Colab would supply the training data for better coding performance. I would argue that Colab is mostly used for non-complex (e.g., small number of variables) and self-contained (i.e., runnable in one page) code that can’t train a model for multi-folder and multi-page projects that rely on global connections, which real-life coding would often require.
at most, the result could be useful for fellow students
I have no doubt other cohorts would rate differently
it is known (on HN at least) e.g. that SWEs tend to prefer brevity, contrary to these students apparently
...and of course this article itself has a bit of AI-ish tone to it.
Fable & Claude Opus 4.x or 5 are terrible to talk to about anything. I gave up on that entirely and just use Anthropic's models for work.
Any less serious technical work I'll use GPT for, as the usage limits are quite fantastic.