Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

84% Positive

Analyzed from 1951 words in the discussion.

Trending Topics

#skills#team#github#https#tools#things#code#more#com#skill

Discussion (39 Comments)Read Original on HackerNews

foundry27•about 1 hour ago
DO NOT INSTALL THIS VIA NPX OR OPEN THIS REPO IN VSCODE. This repo has been infected by malware.

It seems like it was added in commit 74f317d at 11:06 UTC today, with five new hidden files being added under .claude and .vscode that together seem designed to either a) autorun a vscode tasks.json entry, or b) run a Claude session start hook, that will execute a large obfuscated payload. The payload looks like it will fingerprint your system and try to exfil your GitHub tokens.

Edit:

- It also exfils your AWS credentials (~/.aws/credentials, ~/.aws/config), named AWS profiles, and AWS secret managers and SSM parameter store contents

- Same with K8s secrets, with specific searches for GitHub and npm tokens, AWS keys, GCP keys, Azure keys, Stripe keys, Slack tokens, and Twilio keys

- Same with HashiCorp vault contents

- It will try to use your GitHub tokens (if they have the workflow permission) to run actions on your repository and try to exfiltrate secrets from there

- It will try to read a whole bunch of files from your local environment. I didn’t manage to extract the exact file list, unfortunately.

- If the normal C&C server is not available, it tries to create / select a GitHub repo, and commits your data as results-*.json files 100kb at a time

- It also has a bunch of stealth and persistence measures that I’m not qualified to really analyze. Don’t assume that deleting the files is necessarily enough.

Rotate your keys, folks.

zuzululu•36 minutes ago
yeah i had this happen before there was a skill i downloaded and after that i noticed that it was doing strange network behavior

never ever trust a skill especially in an age where its possible to bot submissions on HN (very easy to farm and create voting rings now with AI)

unfortunately this proves that Dang's work has limits, wouldn't be surprised if we've already been seeing manipulation of HN front page for quite some time now

tosh•about 2 hours ago
I know it sounds a bit counter-intuitive but

try a fresh coding session without skills, agents.md, system prompt and additional tools

I think you will be positively surprised how good current models like GPT 5.6 Sol are when they are not oversteered and context spammed

Here is a task (python templating) with 9 runs with OpenCode, Pi and smol

https://smolenv.com/t/nested-template-includes-60636/

you can read each run step by step and see what the agents are doing and how the system prompt and available tools are steering their behaviour to take longer and higher cost

(disclaimer: I'm working on smol)

hansonkd•about 2 hours ago
Also just routine upkeep. Over time if you keep adding new rules these things can get bloated. I deleted 90% of my spec and things improved dramatically.
giancarlostoro•about 1 hour ago
I agree. It's worthwhile to reset your memories and instructions for your models after a major version or two and finetune your needs with the newer model in mind.

You can always backup your old instructions / memories. Personally I try to be hands-off with rules and very simple.

tosh•about 2 hours ago
Pi is pretty good actually

(it has way less context spam, fewer tools and smaller system prompt than the usual suspects)

databricks also looked at this: https://earendil.com/posts/pi-autoresearch-and-databricks/

icedchai•about 1 hour ago
I like the smol philosophy. I've been wondering how much of this stuff is closer to meaningless incantations, voodoo with no real benefit.
tosh•about 2 hours ago
I'm not saying system prompts, agents.md, tools, skills don't have their place

actually the opposite: they matter a lot because they do steer the agent

with great power comes great responsibility

j45•about 2 hours ago
This can work great until the models get tweaked and you realize you’ve been working based on the currents of the model this week.

Creating a robust enough setup that can work on different models or different versions of a model is critical.

tosh•about 1 hour ago
I agree in principle but in practice this is not easy

even with a robust, well thought through setup every model behaves in its own way, some adhere more to a system prompt, another model has more recent cut-off time and knows about new parts in the stdlib

some oversteer, some understeer …

if you want the best performance unfortunately there is not really a way other than to constantly adapt the harness/clutches/context to the model du jour

j45•13 minutes ago
I agree in any practice, it's not easy :)

In a non-deterministic world the compiler is no longer consistent anyways.

cautiouscat•about 3 hours ago
IMO, these agentic guard rails aren't the answer. It seems like we're seeing that the more you stuff context, the more the agents forget and don't follow the guidelines.[1]

Things like ArchUnit, static analyzers, and other deterministic tools can help with lower level things like architecture. For higher up stuff, I am increasingly feeling like agents don't guarantee anything and in many cases its just the opposite. This is where a thoughtful engineer and reviewer can keep things in check.

It's possible I'm off base here, but I can't make heads or tails of the LLM written readme.

1: https://arxiv.org/html/2510.05381v1

jondwillis•about 2 hours ago
I’ve found hooks to be really useful for correcting or inducing certain behaviors/outcomes that tend to occur during agentic development.

For example, I’m working on a project to remake the Final Fantasy XI client. My repository has a bunch of git submodules that reference other peoples’ related efforts, and an open source server. For example, despite including CLAUDE.md to suggest otherwise, Claude Code always ends up writing these insanely dense comments referring to specific files and lines of code in submodules. When those submodules update, now the comments are no longer correct.

So I added a hook to detect when comments are included. Then, for example, I have a deterministic heuristic and script involved to remove some kinds of comments, and another that asks Haiku to quickly LLM-as-a-judge whether or not to edit/remove the comment.

Similarly, I have a stop hook that reminds Claude to commit its code logically on main, noting that other changes may have been added by other concurrent sessions (I avoid worktrees and even branching for this particular project and stage of development.) It works well.

taude•about 2 hours ago
Our skills and agents use deterministic tools for things like: linting, ayy1, etc....
enraged_camel•about 1 hour ago
Yeah. One of the most productive things we did was run a dynamic workflow with Fable to scan for issues in the repo, then have it go through the findings and determine which ones could have been found with deterministic tooling, either tooling that already exists, or could be created from scratch and wired up. The resulting changes we made have been paying off in droves.
kanfilior•about 3 hours ago
see the architect adr and Product pdr, also the whole idea is to have an index of context directive similar to skills which the modal invoke
ctpmpse•about 1 hour ago
Is this commit legit or from today's npm Worm? https://github.com/tikalk/adlc-team-skills/commit/74f317d63a...
jillesvangurp•about 2 hours ago
I created a simple git repository with company skills. Basically just a collection of skills around tools and practices we share. One of the skills is "update company skills" this simply pulls the changes from git and wires them into the user's ~/.codex directory. You can probably do something similar for claude code.

This is far from perfect but we're in this weird transition phase where none of the major AI tool providers are really focusing much on team use of their stuff. But I expect that will start changing soon.

Current tools mostly focus on individuals doing things in isolation. And of course in a team there's more to collaborating than throwing stuff at each other via github. A central repository of company skills is merely our way of improvising a solution.

I find it interesting that Anthropic hired a few of the key people behind Zulip recently. Team chat with tightly integrated AI tools could be a missing piece here. Team communication flows and processes, including ways of working and guardrails are sort of the next piece of the puzzle here. Going from everyone doing their own thing to teams and companies doing things together is going to be a bit of a journey.

taude•about 2 hours ago
We have the exact same thing, we actually install a "plugin" engineering/org wide, and one of the skills in there, is a /skill-mananger skill that can browse other skills in another repo, install, uninstall, and create new ones with all our specific requirments. Semantic versioning allows us to have a hook to auto apply updates, too.

Seems to be helping people share things around the org.

appplication•about 3 hours ago
I see what this is going for but can’t help but feel like it’s overwrought. It feels a bit like the most likely outcomes are increased token burn, review surface, and time per task vs not using this.

Who knows, maybe that’s a good thing.

jlund-molfese•about 3 hours ago
I feel that way about a lot of very long instructions.

It seems that many of these projects are not benchmarked, so it's difficult to know whether there is an improvement in any circumstance, and what the cost is. Of course, a benchmark will be fuzzy, because codebases are all different, but it'd be a start.

mhitza•about 3 hours ago
Relinking the recent study which argues that many instructions in the context are not followed through https://news.ycombinator.com/item?id=49096969

Many of these skills, rules, "playbooks", and such are kitchensink attempts at steering the model. It augments the model to frame it's reasoning according to project rules and flows, but cannot be really trusted to adhere to it. More of a vibe guideline.

kanfilior•about 3 hours ago
see the team-boot skill, it is lean
apt-apt-apt-apt•about 2 hours ago
I feel like this is going to add 100K tokens and 10K rules to everything that the agent is trying to do. Sort of like dumping the Clean Code book into context and saying, hey now you know how to code cleanly, write great code now
zuzululu•about 2 hours ago
Exactly, none of these skills really help other than bloating your token usage

Got rid of superpowers and other useless skills

Your LLM is more than capable of learning from the sea of knowledge

Stitch4223•about 3 hours ago
The promise is nice, yet I find the readme very hard to understand. So I’m not sure how it solves these problems exactly.

First off I would expect a (team) methodology to be referenced. There are tons to choose from. From that point on other terms may make more sense.

For example:

“ Pillar 2: Product Strategy & Architectural Governance (PDRs & ADRs)

Factor III — Mission Definition. Factor IV — Structured Planning. Factor IX — Traceability.”

Why are the Roman numerals in that order, what do they reference and why. Why is this Pilar 2. Etc. It’s easy to get lost in this, even if it would be the best approach in the world.

dawnerd•about 2 hours ago
I couldn’t get past the first couple lines. It’s clear no human reviewed that.
kanfilior•about 3 hours ago
Thanks for the feedback, working on that
hankbond•about 3 hours ago
Write your README or ME won't READ it.
kanfilior•about 3 hours ago
Sure, will be fixing that.
mtzaldo•about 3 hours ago
why is so hard for people to create a CONTRIBUTE.MD file which tells anyone (including AI) how to contribute to the repo. You can also set gates and everything.
dominotw•about 2 hours ago
> Stop Vibe Coding in Silos. Build a Shared Cognitive Layer for Your Engineering Team.

I dont want to be biased but i can get myself to read this after an opening like that .

Advertisement
ppeetteerr•about 2 hours ago
I hate to be that guy, but I'm not reading a readme written by AI, nor using their tool. You don't care to put in a few hours to describe your project, I'm not going to bother learning about it.
xyzsparetimexyz•about 3 hours ago
Slop
clamshelldev•about 2 hours ago
The split that has worked best for me is to keep agent instructions about intent and workflow, then move anything mechanically checkable into tests, linters, or build gates. More rules in context are not enforcement.

For a central skills repository I would also record the exact rules revision in each run or generated artifact. Otherwise a failed run becomes hard to reproduce after the shared repository changes. It would be useful if each skill declared which claims are advisory and which are backed by a command the agent can execute and verify.

kanfilior•about 3 hours ago
I work with an engineering team using Claude Code, Codex, and OpenCode daily. Early on, we created a custom fork of spec-kit to set up our team directives repository and distribute custom agent workflows.

While that fork got us started, maintaining a custom fork of a CLI just to ship prompt workflows created constant merge debt and maintenance headaches. Every session, agents would still drift or forget our architecture decisions, and prompt shortcuts alone couldn't enforce team-wide standards across different developer tools.

To eliminate the fork entirely, we decoupled our workflow skills into adlc-team-skills, built on the open Agent Skills standard (SKILL.md). You install them into any repo with npx skills add tikalk/adlc-team-skills.

The setup works across a few core layers:

On session start, team-boot auto-loads your team constitution from Git and dynamically fetches only the rules, PDRs, and ADRs relevant to the active task — zero prompt-wall bloat.

For product and architecture strategy, product decisions are captured as Product Decision Records (PDRs) and compiled into PRD.md, while architectural decisions use Rozanski and Woods viewpoints composed into AD.md.

For execution, mission-brief acts as an autonomous pipeline runner that derives a formal contract (Goal, Constraints, Non-Goals, Success Criteria) and walks a specify-plan-implement-converge loop. When an agent fails, you edit the spec, not just the code.

In v0.15.0, mission-brief auto-discovers installed skills at runtime. Whether you have spec-kit, OpenSpec, Matt Pocock's skills, Addy Osmani's checklists, or custom skills installed side-by-side, the LLM dynamically decides which skill fits each pipeline step — letting us run upstream spec-kit directly with zero custom fork code.

What doesn't work well yet: our evals suite holdout-split validation is still manual. The architecture skills work, but multi-view DAG orchestration can be slow on very large codebases.

Repos: - Skills: https://github.com/tikalk/adlc-team-skills - Methodology: https://github.com/tikalk/agentic-sdlc-12-factors - CLI: https://github.com/tikalk/adlc-skills-cli

I'm curious — for those of you managing AI coding agents across engineering teams, how are you balancing team standards with the maintenance overhead of custom agent tooling?

senderista•about 2 hours ago
I thought the HN guidelines forbid AI-generated comments?