this post was submitted on 07 Sep 2026
513 points (97.2% liked)

Technology

87897 readers
4447 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
[–] isVeryLoud@lemmy.ca 1 points 9 hours ago (1 children)

How has tool use been for you? I struggled a lot with tool use with Gemma and Qwen, to the point where I needed to build a healing layer.

Regarding the coding harness, I was looking for something CLI-based or JetBrains-based, and I haven't had much luck getting my local llama.cpp models playing ball with OpenCode. They keep losing context and misusing tools.

I'm not too familiar with Apple containers as I'm running a full Linux stack, but I'll give Pi a try, seems interesting! Does it work for coding tasks or is it strictly an "orchestrator"?

[–] percent@infosec.pub 2 points 1 hour ago* (last edited 18 minutes ago) (1 children)

Tool use with Gemma has been hit or miss. I wouldn't rely on it for anything unsupervised.

Tool use for Qwen3.6 has been great lately, but I do remember seeing some issues with it too, a while back. I don't remember when/why the issues cleared up (I have tweaked configs a bit over time), but switching to Pi definitely helped.

I do remember having a lot more problems in OpenCode and it was practically unusable (which is why my recent experience with VS Code was surprising). I'd definitely recommend trying Pi.

A fresh Pi install is very minimal by design. The system prompt is tiny, so it's a pretty good fit for small LLMs like these. It's sort of like Neovim: Nothing fancy out of the box, but you can add lots of fancy things to it. I containerize it because I don't like giving LLMs (especially these small ones) unrestricted access to my host computer – though, I have not seen any signs of it accidentally doing something destructive, which is surprising.

There are similar alternatives to Apple Container for Linux (e.g. Docker Sandboxes, muvm, Firecracker). There's also this thing made specifically for Pi called Gondolin. I haven't tried it yet, but I may end up switching to that if it could simplify my stack.

Here's my current llama-swap/llama.cpp config for Qwen3.6 35B-A3B:

qwen3.6-35b-a3b:
    name: "Qwen3.6 35B-A3B (Coding)"
    proxy: "http://127.0.0.1/:$%7BPORT%7D" # If you're seeing a `/` after `127.0.0.1` here, don't include it. I think something in Lemmy is trying to "sanitize" this input by adding the `/`.
    cmd: |
      llama-server
      --port ${PORT}
      --no-webui
      -hf unsloth/Qwen3.6-35B-A3B-MTP-GGUF:UD-Q4_K_XL
      --jinja
      --parallel 1
      --flash-attn on
      --no-mmproj
      --load-mode none
      --reasoning-preserve
      --ctx-size 190000
      --temp 0.6
      --top-p 0.95
      --top-k 20
      --min-p 0.0
      --presence-penalty 0.0
      --repeat-penalty 1.01

A few notes about this config:

  • Now that I think of it, --reasoning-preserve might be another thing that helped with tool calls.
  • Note the -MTP part of the -hf param. MTP helps speed things up. Here's the Huggingface page for this model
  • You can also omit --no-mmproj if you need vision, but it might mean sacrificing speed or context size, so I usually just enable vision in a separate llama-swap model entry to use as needed.
  • Unsloth recommends --repeat-penalty 1.0, but I saw the LLM enter a thinking loop in VS Code, so I bumped it up just a tiny bit to 1.01. I have since seen it do something that resembled the same thought loop, but it was able to recover on its own. Not sure if it's a coincidence or if 1.01 was actually the solution, so worth some experimentation.
[–] isVeryLoud@lemmy.ca 1 points 1 hour ago (1 children)

Cheers, I'll give this a try!

Regarding containerization, familiarize yourself with Dev Containers, they're super useful for limiting agents to your codebase, with the added bonus that any project you work on comes out of the box with the right version of the tools you need.

[–] percent@infosec.pub 2 points 21 minutes ago (1 children)

Yeah I used to use dev containers. Containers aren't generally a secure sandbox. They're a great guardrail for preventing accidents, but not so much with a malicious prompt injection. (I might be overly paranoid about these things.)

For tool version management, I usually set up a Nix Flake for each project (which also works inside dev containers).

[–] isVeryLoud@lemmy.ca 1 points 16 minutes ago

You're right, they're not watertight. But I'm not trying to defend against MPI, just hallucinations. Never looked into Nix flakes, I always just used OCIs.