為什麼不用現成的 Obsidian plugin?
社群有幾個 Ollama integration plugin,我試過兩個。問題不在功能,在「邊界控制」:
- 多數 plugin 是把整篇筆記丟進 prompt,無法精準選取「我只要摘要這三段」。
- prompt 是寫死在 plugin 設定裡,要改 prompt 就要進設定介面,動線不順。
- 多數 plugin 為了通用性會引入 model selector、temperature slider,這些對日常使用是雜訊。
我要的是:選一段文字 → 跑某個固定任務 → 結果貼回筆記。三件任務各自就是一個 shortcut,不需要選單。
Why not an existing Obsidian plugin?
The community has several Ollama integration plugins. I tried two. The issue isn't features, it's boundary control:
- Most ship "send the whole note as prompt" — no precise selection like "summarize just these three paragraphs."
- Prompts live in plugin settings; tweaking one means digging through a UI.
- Many add a model picker and temperature slider in the name of generality. For daily use that's noise.
What I want: select text → run one fixed task → paste result back. Three tasks, three shortcuts. No menus.
架構:60 行 shell + macOS Shortcuts
Obsidian 選取文字
↓ ⌘⌥1 / ⌘⌥2 / ⌘⌥3
macOS Shortcuts(接住選取的文字)
↓ Run Shell Script
~/bin/llm-helper.sh {task} {selection}
↓ curl
Ollama HTTP API (localhost:11434)
↓ JSON response
回貼到 Obsidian(蓋掉原本的選取)
三個任務(summary / translate / tag)共用一支 shell script,靠第一個參數切分支。macOS Shortcuts 負責「截取選取 → 餵給 script → 把 stdout 蓋回去」。
Architecture: 60 lines of shell + macOS Shortcuts
Obsidian: select text
↓ ⌘⌥1 / ⌘⌥2 / ⌘⌥3
macOS Shortcuts (captures the selection)
↓ Run Shell Script
~/bin/llm-helper.sh {task} {selection}
↓ curl
Ollama HTTP API (localhost:11434)
↓ JSON response
Paste back into Obsidian (replaces the selection)
Three tasks (summary / translate / tag) share one script, branched on the first argument. macOS Shortcuts handles "capture selection → pipe to script → replace selection with stdout."
Shell script
~/bin/llm-helper.sh:
#!/bin/bash
set -euo pipefail
TASK="${1:-}"
INPUT="${2:-}"
MODEL="gemma4:e2b"
ENDPOINT="http://localhost:11434/api/generate"
if [[ -z "$TASK" || -z "$INPUT" ]]; then
echo "usage: llm-helper.sh {summary|translate|tag} <text>" >&2
exit 1
fi
case "$TASK" in
summary)
PROMPT="請務必使用繁體中文台灣用語回答。請用三句話以內摘要下面這段文字,
直接回摘要本體,不要前綴、不要解釋:
---
$INPUT
---"
;;
translate)
PROMPT="請務必使用繁體中文台灣用語翻譯下面這段英文,
保留原文的技術名詞(用括號註原文),直接回譯文,不要解釋:
---
$INPUT
---"
;;
tag)
PROMPT="請務必使用繁體中文台灣用語回答。請為下面這段筆記建議 3-5 個
Obsidian tag(kebab-case,不含 # 字元,用逗號分隔),
不要解釋、不要其他文字、只回 tag 清單:
---
$INPUT
---"
;;
*)
echo "unknown task: $TASK" >&2
exit 2
;;
esac
curl -s "$ENDPOINT" \
-H "Content-Type: application/json" \
-d "$(jq -nc \
--arg model "$MODEL" \
--arg prompt "$PROMPT" \
'{model: $model, prompt: $prompt, stream: false, think: false, options: {temperature: 0.2}}')" \
| jq -r '.response' \
| sed -e 's/^[[:space:]]*//' -e 's/[[:space:]]*$//'
五件事值得指出:
- prompt 都繁中前綴——
gemma4:e2b預設英文,這句不能省(來源)。 stream: false——shortcut 需要一次拿到完整回應,不需要 streaming。temperature: 0.2——摘要、翻譯、tag 都要穩定,不要創意。think: false——gemma4:e2b是 thinking 模型,沒有這個 flag,num_predict配額會先被內部推理過程吃掉,response欄位回傳空字串。這個坑踩過一次就不會忘。jq -nc組 JSON——避免 shell 變數插值踩到引號跳脫。
Shell script
~/bin/llm-helper.sh:
#!/bin/bash
set -euo pipefail
TASK="${1:-}"
INPUT="${2:-}"
MODEL="gemma4:e2b"
ENDPOINT="http://localhost:11434/api/generate"
if [[ -z "$TASK" || -z "$INPUT" ]]; then
echo "usage: llm-helper.sh {summary|translate|tag} <text>" >&2
exit 1
fi
case "$TASK" in
summary)
PROMPT="請務必使用繁體中文台灣用語回答。請用三句話以內摘要下面這段文字,
直接回摘要本體,不要前綴、不要解釋:
---
$INPUT
---"
;;
translate)
PROMPT="請務必使用繁體中文台灣用語翻譯下面這段英文,
保留原文的技術名詞(用括號註原文),直接回譯文,不要解釋:
---
$INPUT
---"
;;
tag)
PROMPT="請務必使用繁體中文台灣用語回答。請為下面這段筆記建議 3-5 個
Obsidian tag(kebab-case,不含 # 字元,用逗號分隔),
不要解釋、不要其他文字、只回 tag 清單:
---
$INPUT
---"
;;
*)
echo "unknown task: $TASK" >&2
exit 2
;;
esac
curl -s "$ENDPOINT" \
-H "Content-Type: application/json" \
-d "$(jq -nc \
--arg model "$MODEL" \
--arg prompt "$PROMPT" \
'{model: $model, prompt: $prompt, stream: false, think: false, options: {temperature: 0.2}}')" \
| jq -r '.response' \
| sed -e 's/^[[:space:]]*//' -e 's/[[:space:]]*$//'
Five things worth pointing out:
- All prompts have the Chinese prefix.
gemma4:e2bdefaults to English (prior article). Don't drop the line. stream: false— the Shortcut needs a complete response in one shot.temperature: 0.2— summary, translation, tags all want stability, not creativity.think: false—gemma4:e2bis a thinking model. Without this flag, thenum_predictbudget gets eaten by the internal reasoning pass first, andresponsecomes back empty. A lesson you only need to learn once.jq -ncto build the JSON — avoids shell-variable quoting hell.
macOS Shortcuts 設定
三個 shortcut,邏輯相同,差別只在參數:
- 新建 Shortcut:「Quick Action」類型
- 接收:在 Obsidian 中接收「Text」
- 動作 1:Get Text from Shortcut Input
- 動作 2:Run Shell Script
- Shell: /bin/zsh - Pass input: as arguments - Script: $HOME/bin/llm-helper.sh summary "$1"
- 動作 3:Replace Selection(用 stdout 蓋回去)
三個 shortcut 分別命名 Obsidian Summary、Obsidian Translate、Obsidian Tag,到「System Settings → Keyboard → Keyboard Shortcuts → Services」綁 ⌘⌥1 / ⌘⌥2 / ⌘⌥3。
macOS Shortcuts setup
Three Shortcuts, same shape, different argument:
- New Shortcut → Quick Action
- Receives: Text from Obsidian
- Action 1: Get Text from Shortcut Input
- Action 2: Run Shell Script
- Shell: /bin/zsh - Pass input: as arguments - Script: $HOME/bin/llm-helper.sh summary "$1"
- Action 3: Replace Selection (with stdout)
Name them Obsidian Summary, Obsidian Translate, Obsidian Tag, then bind ⌘⌥1 / ⌘⌥2 / ⌘⌥3 under System Settings → Keyboard → Keyboard Shortcuts → Services.
三個任務實測
我從自己的 Obsidian vault(這個 repo 的 content/ 資料夾本身就兼作 vault)為每個任務挑了三種不同類型的輸入,各跑 10 次取平均,三個任務共 90 次真實 Ollama 呼叫,零空回應。原始記錄在 tests/outputs/functional_test_results.jsonl,彙總在 tests/outputs/functional_test_summary.json。
| 任務 | 平均推論時間 | 平均 tokens | 主觀品質 (1-5) | 我會不會用 |
|---|---|---|---|---|
| Summary | 1.6 秒 | 61 | ||
| Translate | 1.5 秒 | 56 | ||
| Tag | 1.1 秒 | 34 |
推論時間比原本預估的 8–20 秒快很多——gemma4:e2b 只有 5.1B,這三個任務的輸出也都不長。倒是 tag 任務格式穩定度不如預期:90 次裡有好幾次吐出 # 開頭或英文詞,跟下面「限制與下一步」段的觀察一致。
Summary 樣本
輸入(節錄自我自己的 ARCHITECTURE.md 一段):
此監看倉並非強迫所有站點使用同一個 SSG、theme 或 deploy 目標。
每個站點保留最適合自身內容的技術棧。共享的是 knowledge/ skills/
pipeline/ metrics/ n8n-workflows/;不共享 SSG 選擇、視覺識別、
部署目標、是否雙語。
輸出(Gemma 4 e2b):
各站點可根據自身內容選擇最適合的技術棧,不強制使用統一的SSG、theme或部署目標。
僅共享知識、技能、管道、指標等資訊。SSG選擇、視覺識別、部署目標及雙語選項不共享。
Translate 樣本
輸入:The unified memory architecture on Apple Silicon allows large models to run without dedicated VRAM, but inference throughput remains bound by memory bandwidth, not compute.
輸出:
Apple Silicon 上的統一記憶體架構 (unified memory architecture) 允許大型模型在沒有專用
VRAM 的情況下運行,但推論吞吐量 (inference throughput) 仍然受限於記憶體頻寬
(memory bandwidth),而非運算能力 (compute)。
Tag 樣本
輸入(一段關於 RAG chatbot embedding retrieval 的筆記):
今天研究了怎麼幫 RAG chatbot 做 embedding retrieval。用 nomic-embed-text 把每篇文件
切成 300-500 字的 chunk,算 cosine similarity 找出最相關的前 5 段,再把這些段落塞進
prompt context 讓 LLM 生成回答。測試發現 chunk 太短會失去上下文,chunk 太長又容易稀
釋掉關鍵句子的權重,抓 top-k 的 k 值也要跟 chunk 大小一起調。晚點想試試看 hybrid
search,把關鍵字比對(BM25)跟向量相似度加權混合,看看能不能提升召回率。
輸出:
obsidian tag:
RAG
embedding retrieval
nomic-embed-text
chunking
context window
hybrid search
bm25
這次輸出正好示範了 tag 任務的問題——prompt 明明要求「不要解釋、不要其他文字、只回 tag 清單」,模型還是加了 obsidian tag: 前綴,而且是換行列點而不是逗號分隔,也沒有轉成 kebab-case。90 次功能測試裡這種格式偏移不是單一個案。
Three tasks, measured
I picked three different input types per task from my own Obsidian vault (this repo's content/ folder doubles as the vault), ran 10 trials per type — 90 real Ollama calls across the three tasks, zero empty responses. Raw records: tests/outputs/functional_test_results.jsonl; aggregates: tests/outputs/functional_test_summary.json.
| Task | Avg inference time | Avg tokens | Subjective quality (1-5) | Do I keep using it? |
|---|---|---|---|---|
| Summary | 1.6s | 61 | ||
| Translate | 1.5s | 56 | ||
| Tag | 1.1s | 34 |
Much faster than the 8–20s I'd originally guessed — gemma4:e2b is only 5.1B and none of these three tasks need a long output. Tag format compliance was the shakier one: several of the 90 runs came back with a leading # or a stray English word, matching the observation in "Limits and what's next" below.
Summary sample
Input (excerpt from my own ARCHITECTURE.md):
此監看倉並非強迫所有站點使用同一個 SSG、theme 或 deploy 目標。
每個站點保留最適合自身內容的技術棧。共享的是 knowledge/ skills/
pipeline/ metrics/ n8n-workflows/;不共享 SSG 選擇、視覺識別、
部署目標、是否雙語。
Output (gemma4:e2b):
各站點可根據自身內容選擇最適合的技術棧,不強制使用統一的SSG、theme或部署目標。
僅共享知識、技能、管道、指標等資訊。SSG選擇、視覺識別、部署目標及雙語選項不共享。
Translate sample
Input: The unified memory architecture on Apple Silicon allows large models to run without dedicated VRAM, but inference throughput remains bound by memory bandwidth, not compute.
Output:
Apple Silicon 上的統一記憶體架構 (unified memory architecture) 允許大型模型在沒有專用
VRAM 的情況下運行,但推論吞吐量 (inference throughput) 仍然受限於記憶體頻寬
(memory bandwidth),而非運算能力 (compute)。
Tag sample
Input (a note on RAG chatbot embedding retrieval):
今天研究了怎麼幫 RAG chatbot 做 embedding retrieval。用 nomic-embed-text 把每篇文件
切成 300-500 字的 chunk,算 cosine similarity 找出最相關的前 5 段,再把這些段落塞進
prompt context 讓 LLM 生成回答。測試發現 chunk 太短會失去上下文,chunk 太長又容易稀
釋掉關鍵句子的權重,抓 top-k 的 k 值也要跟 chunk 大小一起調。晚點想試試看 hybrid
search,把關鍵字比對(BM25)跟向量相似度加權混合,看看能不能提升召回率。
Output:
obsidian tag:
RAG
embedding retrieval
nomic-embed-text
chunking
context window
hybrid search
bm25
This run is a good demonstration of the tag task's problem — the prompt explicitly says "no explanation, no other text, just the tag list," and the model still prepended obsidian tag:, one-per-line instead of comma-separated, and never converted to kebab-case. This format drift wasn't a one-off across the 90 functional-test runs.
跟原本用 Claude 比
| 指標 | Claude(之前) | Gemma 4 e2b(現在) |
|---|---|---|
| 每天 prompt 數 | ~30 | |
| 每次延遲 | 1-3 秒 | |
| 摘要品質 | 5/5 | |
| 翻譯品質 | 5/5 | |
| Tag 品質 | 4/5(有時太發散) | |
| 月成本(推估) | Claude Pro 部分使用量 | NT$ 電費 |
| 私隱 | 上雲 | 本機 |
我的判斷不是「Gemma e2b 全面取代 Claude」,是把這三個重複任務從 frontier 模型卸下來——Claude 額度留給真的需要它的事(例如寫這篇文章的英文版)。
Compared to my previous Claude flow
| Metric | Claude (before) | gemma4:e2b (now) |
|---|---|---|
| Daily prompts | ~30 | |
| Per-prompt latency | 1-3 s | |
| Summary quality | 5/5 | |
| Translation quality | 5/5 | |
| Tag quality | 4/5 (sometimes too broad) | |
| Monthly cost (est.) | portion of Claude Pro | NT$ electricity |
| Privacy | cloud | local |
The point isn't "Gemma e2b replaces Claude." It's offloading these three repeated chores so frontier-model budget goes to work that actually needs it (like writing the English version of this article).
限制與下一步
- 長文輸入:超過 ~3K token 的筆記,summary 開始忽略後半段。要靠先手動分段。
- Tag 中英混雜:偶爾會吐出英文 tag。在 prompt 加「只用繁中名詞」之後改善但沒完全解決。
- macOS Shortcuts 偶爾搶不到選取:Obsidian 的 markdown 預覽模式選取行為跟原始模式不同,這個我還在繞。
下一步打算把這套擴成第四個任務:「給定一段筆記,找出 vault 內可能相關的其他筆記」——這要接 embedding,會比現在這套複雜,會單獨寫一篇。
Limits and what's next
- Long input: past ~3K tokens, summary starts dropping the back half. Manual chunking is the current workaround.
- Mixed-language tags: occasional English tag slips in. The "Chinese nouns only" instruction helps, doesn't fully fix.
- Shortcut selection misses: Obsidian's preview-mode selection behaves differently from source mode. Still working around this.
Next: extending this to a fourth task — "given this note, find related notes in the vault." That needs embeddings and gets meaningfully more complex; it'll be its own article.
常見問題
Q:為什麼不用 e2b 直接接 Obsidian 的 Smart Connections plugin? A:Smart Connections 是 embedding 為主,這篇做的是 generation。兩件事互補,不互斥。
Q:可以用更大的模型換掉 e2b 嗎? A:架構上完全可以——llm-helper.sh 的 MODEL 那一行改掉就好,其他都不用動。這台機器目前只裝了 gemma4:e2b(5.1B),沒有裝更大量級的版本,所以沒辦法在這裡給你「換更大模型後 tag 一次過率會不會提高」的實測數字。這是誠實的空白,不是我懶得測。
Q:Windows / Linux 怎麼複製這套? A:Shell script 直接通用。macOS Shortcuts 在 Windows 上對應 AutoHotkey、Linux 對應 xdotool + keybinding。邏輯相同。
FAQ
Q: Why not use Obsidian's Smart Connections plugin with e2b? A: Smart Connections is embedding-based. This article is generation. Complementary, not competing.
Q: Can I swap in a bigger model? A: Architecturally, yes — just change the MODEL line in llm-helper.sh. This machine only has gemma4:e2b (5.1B) installed; I haven't run a larger size, so I can't hand you real numbers on whether tag compliance improves with more parameters. That's an honest gap, not something I'm skipping out of laziness.
Q: Windows / Linux? A: The shell script ports directly. Replace macOS Shortcuts with AutoHotkey (Windows) or xdotool + a keybinding (Linux). Same shape.
共振站結論
波形評級:STEADY_PULSE
適合你,如果:
- 用 Obsidian 寫筆記,且筆記量到了一定規模
- 已經有 ollama 在跑、或願意花 30 分鐘設好
- 想把 frontier 模型額度省給真正需要的事
- 在意私隱:筆記內容不上雲
暫不適合,如果:
- 你的筆記長度普遍超過 5K token(要等 streaming / 分段方案)
- 不想碰 shell script 與 macOS Shortcuts
Verdict
Waveform: STEADY_PULSE
Good fit if you:
- Use Obsidian and have enough notes for the chores to recur daily
- Have ollama already running, or are willing to spend 30 minutes setting it up
- Want to free frontier-model budget for work that actually needs it
- Care about privacy: notes never leave the machine
Skip if you:
- Notes routinely exceed 5K tokens (wait for the streaming/chunking version)
- Don't want to touch shell scripts or macOS Shortcuts