Short answer: the reply stopped in the middle of a sentence because the model hit the output limit. The response said finishReason: MAX_TOKENS. In my run the model used 573 tokens for thinking and only 23 for the visible text, and I had set maxOutputTokens to 600. I raised the limit to 2048 and the replies came out complete.
What I saw
My n8n workflow asks Gemini to draft a short reply to a blog comment. One reply ended like this, with no final punctuation:
Hello! To fix a 404 error on your Production URL, first make sure that your workflow is published
The HTTP Request node did not fail, and the response looked normal. The cut-off showed only in the text.
What the response said
Here are the relevant fields of the response. I removed a long signature string and the response ID.
{
"finishReason": "MAX_TOKENS",
"usageMetadata": {
"promptTokenCount": 1271,
"candidatesTokenCount": 23,
"thoughtsTokenCount": 573,
"totalTokenCount": 1867
},
"modelVersion": "gemini-3.5-flash"
}
Two things stand out:
finishReasonisMAX_TOKENS, notSTOP. A normal ending showsSTOP.- The model spent 573 tokens on thinking (
thoughtsTokenCount) and 23 on the visible answer (candidatesTokenCount).
The numbers
| Value | Tokens |
|---|---|
Prompt (promptTokenCount) | 1271 |
Visible answer (candidatesTokenCount) | 23 |
Thinking (thoughtsTokenCount) | 573 |
Total (totalTokenCount) | 1867 |
The total matches 1271 + 23 + 573. Thinking plus answer is 573 + 23 = 596 tokens, which is almost exactly the 600 I had set as maxOutputTokens. In this response the thinking tokens used most of the limit, so only a short part of the answer fit. I did not look for a statement about this in the documentation. Treat it as what this one response shows.
How I fixed it
I raised the output limit in the request body. This is the setting I use now:
generationConfig: { temperature: 0.3, maxOutputTokens: 2048 }
After that, the replies I tested ended with a full sentence and a period. For example, a later reply was 283 characters long and ended normally.
How to detect it before you publish
Do not rely on the text alone. Read finishReason in a Code node after the HTTP Request node and treat anything other than STOP as unfinished. Here is the check I use, where r is the HTTP Request output:
const cand = r && r.candidates && r.candidates[0];
const finish = cand && cand.finishReason;
const truncated = !!finish && finish !== 'STOP';
In my workflow, a truncated reply adds a warning to the Telegram message that asks me to approve the reply, so I can discard it instead of publishing it.
Environment
- n8n 2.8.3, self-hosted
- HTTP Request node calling the Gemini API
- Model: gemini-3.5-flash (the value of
modelVersionin the response)
What I did not test
- Other models
- Other limit values between 600 and 2048
- Settings that reduce thinking, such as a thinking configuration
- Other n8n versions
- What the documentation says about how thinking tokens count against the limit
Pingback: n8n으로 워드프레스 댓글에 AI 답글 초안을 만들고 텔레그램에서 승인한 뒤 게시하기 - AI Brief