Gemini Reply Cut Off Mid-Sentence in n8n: finishReason MAX_TOKENS and maxOutputTokens

Short answer: the reply stopped in the middle of a sentence because the model hit the output limit. The response said finishReason: MAX_TOKENS. In my run the model used 573 tokens for thinking and only 23 for the visible text, and I had set maxOutputTokens to 600. I raised the limit to 2048 and the replies came out complete.

What I saw

My n8n workflow asks Gemini to draft a short reply to a blog comment. One reply ended like this, with no final punctuation:

Hello! To fix a 404 error on your Production URL, first make sure that your workflow is published

The HTTP Request node did not fail, and the response looked normal. The cut-off showed only in the text.

What the response said

Here are the relevant fields of the response. I removed a long signature string and the response ID.

{
  "finishReason": "MAX_TOKENS",
  "usageMetadata": {
    "promptTokenCount": 1271,
    "candidatesTokenCount": 23,
    "thoughtsTokenCount": 573,
    "totalTokenCount": 1867
  },
  "modelVersion": "gemini-3.5-flash"
}

Two things stand out:

  • finishReason is MAX_TOKENS, not STOP. A normal ending shows STOP.
  • The model spent 573 tokens on thinking (thoughtsTokenCount) and 23 on the visible answer (candidatesTokenCount).

The numbers

ValueTokens
Prompt (promptTokenCount)1271
Visible answer (candidatesTokenCount)23
Thinking (thoughtsTokenCount)573
Total (totalTokenCount)1867

The total matches 1271 + 23 + 573. Thinking plus answer is 573 + 23 = 596 tokens, which is almost exactly the 600 I had set as maxOutputTokens. In this response the thinking tokens used most of the limit, so only a short part of the answer fit. I did not look for a statement about this in the documentation. Treat it as what this one response shows.

How I fixed it

I raised the output limit in the request body. This is the setting I use now:

generationConfig: { temperature: 0.3, maxOutputTokens: 2048 }

After that, the replies I tested ended with a full sentence and a period. For example, a later reply was 283 characters long and ended normally.

How to detect it before you publish

Do not rely on the text alone. Read finishReason in a Code node after the HTTP Request node and treat anything other than STOP as unfinished. Here is the check I use, where r is the HTTP Request output:

const cand = r && r.candidates && r.candidates[0];
const finish = cand && cand.finishReason;
const truncated = !!finish && finish !== 'STOP';

In my workflow, a truncated reply adds a warning to the Telegram message that asks me to approve the reply, so I can discard it instead of publishing it.

Environment

  • n8n 2.8.3, self-hosted
  • HTTP Request node calling the Gemini API
  • Model: gemini-3.5-flash (the value of modelVersion in the response)

What I did not test

  • Other models
  • Other limit values between 600 and 2048
  • Settings that reduce thinking, such as a thinking configuration
  • Other n8n versions
  • What the documentation says about how thinking tokens count against the limit

Related

1 thought on “Gemini Reply Cut Off Mid-Sentence in n8n: finishReason MAX_TOKENS and maxOutputTokens”

  1. Pingback: n8n으로 워드프레스 댓글에 AI 답글 초안을 만들고 텔레그램에서 승인한 뒤 게시하기 - AI Brief

Leave a Reply

Discover more from AI Brief

Subscribe now to keep reading and get access to the full archive.

Continue reading