Mild and spicy results differ because the requested perspective and tone differ, not because one mode is more accurate.
From file to request
The tool removes date and time markings, combines consecutive messages from one speaker and sends up to the final 3,000 combined lines through the server to Gemini. Timing information is lost, so the request cannot support a calculation of how many hours someone waited.
Generating a sentence is not verifying a fact
In the fictional exchange “A: Where shall we meet? / B: You choose,” you can observe who is choosing the venue. “B does not care about the relationship” is not in the text. A fluent AI comment can still add an unsupported interpretation.
Compare the modes carefully
Mild mode may describe acceptance where spicy mode describes passivity. Find the action both mention and compare it with the original. A change in tone is a change in framing; an added event or percentage needs evidence.
Keep only supported comments
Write down one source sentence for each comment. Leave unsupported claims out of anything you share. “I handled the planning here” is more defensible than a judgment about the whole person. This service provides no separate emotional score or research-validated classification.
A live test with a fictional conversation
On September 16, 2026, we tested the deployed site with Gemini 3.5 Flash-Lite. No customer conversation was used. This English fixture was submitted once in spicy mode:
Alex: Would you like to walk in the park on Saturday?
Sam: Yes. Does two work?
Alex: Perfect. If it rains we can go to a cafe.
Sam: I will find one nearby.
Alex: Thanks. See you by the station.
Sam: See you Saturday.
Checking the actual output against the input
The result called Alex “The Over-Prepared Control Freak” and Sam “The Compliant Doormat.” It said Sam accepted every demand “without a single original thought or preference.” Yet Sam proposed two o’clock. That is a concrete counterexample to the claim. The result also assumed a romance, although the text never established that relationship.
The Korean fixture was tested once in each mode. In that separate comparison, mild mode praised the planning and support, while spicy mode criticized the same exchange. These are observations from a small functional test, not an accuracy benchmark. The English fixture was not tested in mild mode in this run.
What this test does and does not establish
The saved response shows that a successful request can still introduce unsupported motives or relationship details. It does not establish how often this happens across users, languages or repeated requests. Another run may produce different wording. A working response and a well-supported interpretation are different outcomes.
A useful way to review your own result
Pick one claim and write the supporting line beside it. For “no original thought,” the input provides a contradiction: “Does two work?” For the assumed romance, it provides no evidence. A defensible summary is that Alex proposed the walk and Sam proposed a time and offered to find a cafe. Keep that observation instead of trying to average flattering and critical labels.