What a hallucination looks like in a Microsoft 365 tenant
A large language model predicts likely words. It does not look facts up the way a database does. Microsoft's own transparency material for Microsoft 365 Copilot describes the failure as ungrounded content: a response that is not supported by the material the model was given8. The public-service guidance from the Government Chief Digital Officer uses the plainer word hallucination, and treats it as a known property of generative tools rather than a rare bug4. Microsoft has also renamed the product itself: Microsoft 365 Copilot is now called Microsoft Copilot, and the checks in this guide apply under either name8.
Copilot is less prone to invention than a bare chatbot because it grounds many answers. When a person signs in with a work account and asks about their own files, Copilot retrieves emails, chats and documents they already have permission to open, and builds the answer from them6,7. Grounding narrows the problem. It does not remove it. In a working tenant, most errors fall into a handful of patterns:
- Wrong source: Copilot finds last year's price list or a superseded policy and summarises it accurately.
- Blended sources: two clients, two projects or two versions of a contract are merged into one tidy paragraph.
- Missing context: a meeting recap leaves out the objection that changed the decision.
- Invented detail: a date, a clause number, a case name or a figure appears that is not in any source.
- Plausible arithmetic: an Excel summary or a percentage that reads well but was never calculated from the cells.
Why the legal duty sits with the business, not the tool
Under the Privacy Act 2020, information privacy principle 8 says an agency must not use or disclose personal information without taking reasonable steps to make sure it is accurate, up to date, complete, relevant and not misleading1. A Copilot summary of a customer complaint, a staff performance note drafted from Teams chats, or a reference letter built from old emails are all personal information once you save or send them. The Privacy Commissioner expects a person to review generative output before an agency acts on it, precisely because the tools can be confidently wrong2,3.
Accuracy duties do not stop at privacy. The Fair Trading Act 1986 prohibits misleading or deceptive conduct in trade, and that covers a product claim on your website whether a person or a model wrote it11. The Courts of New Zealand guidelines for lawyers tell them to check the accuracy of anything a generative AI chatbot provides, including legal citations, and the guidelines for non-lawyers warn that chatbots can make up cases, citations and quotes10,14. MBIE's responsible AI guidance for businesses frames the same idea in management terms: keep a human accountable for what an AI system produces, and match the level of checking to the level of risk5.
Sort the work by what a mistake would cost
Checking every Copilot answer to the same depth wastes the time the tool was meant to save. A better approach is three tiers, written down and agreed with staff. A small Tauranga property management office, for example, could fit its tiers on one page and pin it in the Teams channel where staff share prompts.
- Tier 1, internal and reversible: a first draft of an internal email, a brainstorm, a rewrite for tone. Read it once before you send it. No source check is needed.
- Tier 2, leaves the building or informs a decision: a client letter, a quote cover note, a board paper, a meeting summary circulated to people who were not there. Open every cited source and check names, dates, amounts and commitments against it.
- Tier 3, legal, financial, health or employment consequences: a tenancy notice, a tax position, a disciplinary letter, a safety procedure. Copilot may help with structure and wording, but a qualified person writes or approves the substance from primary documents, not from the summary.
Use the citations Copilot gives you
When Copilot Chat or Copilot in an app draws on a file, an email or a web page, the response carries numbered references. Hover over or select a reference to see the source, and open it. Microsoft's guidance on grounding explains that work answers come from content the signed-in person can already reach, so a reference you cannot open is itself a warning sign worth reporting to whoever runs your tenant6.
Three habits make citations useful rather than decorative. First, check that the reference actually says what the sentence beside it claims; a real citation attached to an invented detail is the most convincing kind of error. Second, look at the date and version of the cited file. Third, notice any sentence with no reference at all. In a work-grounded answer, an unreferenced claim may have come from the model's general training rather than your documents.
If your tenant allows web search in Copilot, answers can also cite public web pages. Treat those as you would any search result: check the publisher and date before you repeat the claim, and prefer the primary source, such as legislation.govt.nz for an Act or the regulator's own site for guidance.
Quick app-by-app checks
Each Microsoft 365 app fails in its own way, so each needs its own short routine. The routines below are short enough to become habit.
- Word: when you ask Copilot to draft from a file, name the file explicitly with a slash reference rather than letting Copilot choose. Then compare every number, date and party name in the draft with that file. Use Track Changes for Copilot rewrites of contracts or policies so a reviewer can see what moved.
- Excel: Copilot uses Excel tools such as tables and PivotTables9, so give it a tidy table with clear headers. When it proposes a formula or a summary, ask it to add the formula as a new column rather than paste values, then spot-check three rows by hand, including the first, the last and one odd one. Pivot the same data yourself once to confirm a total.
- Outlook: before sending a Copilot reply, check who it is addressed to and whether it promises anything, such as a delivery date, a refund or a discount. Summaries of long threads often drop the latest message, so read the most recent email yourself.
- Teams meetings: a recap is only as good as the transcript behind it, and after the meeting Copilot answers questions from the most recent available transcript13. Check the action items with their owners before you circulate them, and correct any names the transcript has got wrong. If a decision matters, confirm it in writing with the people who made it.
- PowerPoint: Copilot will happily generate statistics for a slide. Delete any figure that does not come from a source you can name, or replace it with a sourced one.
Prompts that reduce errors before they happen
The way a request is phrased changes how often Copilot guesses. Give it the source, the scope and permission to say it does not know. A prompt such as summarise the payment terms in the attached supplier agreement, quote the clause numbers, and say not stated if a term is missing will produce fewer inventions than a vague request to summarise the contract.
Ask for structure that makes checking easy: a table with a source column, quotes alongside paraphrases, or a list of assumptions at the end. For anything numerical, ask Copilot to show the calculation. When it refuses or hedges, take that as useful information rather than pushing it until it produces an answer.
- Point to the exact file, email thread or meeting, rather than asking in general.
- Tell it what to do when information is missing.
- Ask for quotes or clause references, not only a paraphrase.
- Keep one topic per chat so earlier material does not leak into a new answer.
A hypothetical example from a small NZ firm
Picture an 8-person Christchurch engineering consultancy preparing a fee proposal for a council job. The project lead asks Copilot in Word to draft the methodology section from two past proposals and the council's request for proposal. The draft is good. It is also wrong in two places: it promises a site visit schedule copied from an older job, and it refers to a standard the council did not mention.
The lead catches both by doing three things. They open each reference and confirm the source file is the current one. They search the request for proposal for every standard named in the draft. They ask Copilot, in the same chat, to list every commitment the draft makes about time, cost or deliverables, and they check that list against the brief. The check is short, and it turns a risky draft into a usable one.
The same firm then adds one line to its quality procedure: any Copilot-assisted document that goes to a client is reviewed by someone who did not write the prompt. That one rule does more than any amount of general warning about AI.
Make checking a team habit, not a personal virtue
Individual vigilance fades after the first month. Build the checks into how work moves instead. NCSC's joint guidance says staff should be trained on how far AI output can be relied on, and on the organisation's process for validating it12. The GCDO guidance adds that accountable people should stay involved in decisions that use AI output4.
Start with a short list of the documents your business would never want wrong: tax returns, employment letters, safety documents, client advice and anything filed with a regulator. Decide who reviews each one when Copilot was involved. Add a simple marker, such as a line in the file properties or a tag in the document library, so a reviewer knows a draft had AI help. Collect real examples of errors your team catches and share them at a monthly team meeting; nothing teaches checking faster than a near miss from your own work.
- Write the three tiers down and pin them where staff work.
- Name a reviewer for each high-risk document type.
- Keep a shared log of caught errors and what would have prevented them.
- Review the rules when Microsoft adds a new Copilot feature to your tenant.