Start your 3-day free trial
Sign up to experience all premium features at no cost.
*Available only to new users. Each user is limited to one trial.


To build website information architecture with AI, begin with evidence of what people actually need to do, not a blank sitemap prompt. Freeze a content inventory and a register of user tasks, then let AI propose groupings, hierarchy, labels, and alternative paths. Treat every proposal as a hypothesis. Card sorting, tree testing, task walkthroughs, accessibility review, and accountable owners determine what is adopted.
The general AI workflow explains how to constrain evidence and verify output. This guide applies it to information architecture, or IA: the organization, naming, and pathways that help people find and understand website content. It does not allow AI to invent user needs or declare an untested navigation successful.
Key Takeaways
- Derive tasks from research, search logs, support questions, and observed behavior.
- Keep tasks, content, pages, features, labels, and navigation components distinct.
- Ask AI for several candidate structures and cite the evidence behind each choice.
- Preserve conflicting mental models rather than averaging them away.
- Test findability with representative participants and realistic tasks.
- Version the approved IA and retain unresolved evidence and exceptions.
A useful project produces a decision trail, not merely a diagram. Its core artifacts are a frozen content inventory, an evidence register, normalized task statements, candidate structures, a label glossary, test protocols and results, an exception log, and an approved version of the hierarchy.
Keep these concepts separate:
| Concept | Example | Why it matters |
|---|---|---|
| User task | Replace an expired payment method | Describes an outcome in the user's language |
| Content | Eligibility rules and instructions | Information needed to complete the task |
| Page | Billing settings help page | One possible delivery container |
| Feature | Update-card form | Interactive capability, not a category |
| Label | Payments and billing | A signpost that must be understood |
| Navigation path | Account → Payments → Update card | One route, not proof that the task is easy |
An organization chart is not automatically an IA. Neither is a database schema, a product team's ownership map, or a list of current URLs. Those inputs may explain constraints, but the structure should support the language and relationships users understand.
Write a project header with the site or section, audience, supported locales, business and user outcomes, research period, analytics period, content snapshot date, known technical constraints, approvers, and excluded areas. Record whether the work is a diagnosis, partial redesign, migration, or new structure.
Inventory every in-scope URL or content object. Include title, purpose, owner, format, audience, locale, lifecycle status, inbound links, traffic or search evidence when approved, and duplication or quality notes. Do not delete apparently unused content during discovery; mark it for a separate governance decision.
Freeze the inventory before asking AI to reorganize it. Otherwise new pages, redirects, or editorial changes can make candidate structures incomparable. If the inventory is incomplete, label the gap rather than allowing the model to fill it from general web conventions.
GOV.UK advises teams to start by learning user needs rather than assuming them.[1] Translate each research observation into a traceable task record, but preserve the original evidence beside the normalized wording.
Useful evidence can include moderated interviews, contextual inquiry, usability sessions, on-site search terms, support tickets, call reasons, feedback, failed journeys, accessibility research, and observed workarounds. A user interview synthesis workflow can help organize notes, but only approved research belongs in the register.
For each task, record:
Do not turn a stakeholder request such as “promote premium services” into a user need without evidence. Do not treat a search query as a complete task without context. Do not let AI fabricate a persona, frequency, motivation, or pain point.
Users often enter a site partway through a larger problem and continue elsewhere. GOV.UK's guidance on mapping a user's whole problem encourages teams to understand the steps before, during, and after their service interaction.[3] Record dependencies and adjacent tasks without claiming the website owns all of them.
A customer journey map can show sequence and handoffs, while IA describes grouping and findability. Do not convert journey stages directly into top-level navigation. “Discover, decide, buy, use” may be a useful analysis frame yet a poor set of labels for someone who wants to replace a lost card.
Also separate task statements from implementation requirements. User stories and acceptance criteria describe product behavior; they do not automatically determine page boundaries or menu levels.
Provide the frozen task register, approved inventory fields, constraints, and an output schema. Remove personal information from research excerpts and use task IDs in the prompt. A bounded instruction can say:
Using only the supplied task and content records, propose three alternative
groupings. For every group and label, cite supporting task IDs. Preserve
conflicting patterns and mark unsupported placement as unresolved. Do not add
users, needs, pages, features, research findings, or success claims. Return a
hierarchy, cross-links, label alternatives, and questions for researchers.
Ask for meaningfully different candidates: perhaps task-oriented, object-oriented, and lifecycle-oriented. Require the model to explain tradeoffs and identify items that need multiple entrances. Reject invented evidence, uncited placements, duplicated IDs, missing inventory items, and a hierarchy deeper than the agreed limit.
NIST's generative AI profile emphasizes lifecycle risk management and evaluation.[4] Here, that means recording the input and model configuration, validating the output schema, testing proposals with people, and monitoring the approved structure after launch.
Review candidate structures before user testing. Remove internal department names unless research shows users understand and seek them. Check whether siblings are conceptually parallel, whether labels overlap, and whether one category has become a miscellaneous drawer.
GOV.UK's service-scoping guidance frames services around what users are trying to achieve rather than existing organizational boundaries.[2] Apply that principle without pretending every website is a government service: define the task and its boundaries in the user's terms, then record organizational constraints separately.
For each proposed label, keep a small glossary with definition, included and excluded content, tested alternatives, locale-specific wording, and owner. The localization glossary workflow is useful when labels span languages. A literal translation may merge distinctions or create new ambiguity, so approve labels per locale.
Card sorting helps investigate how participants group representative items and what they call those groups. Use open sorting to discover patterns, closed sorting to evaluate defined categories, or a hybrid when both questions matter. Choose cards from real tasks and content; avoid internal jargon that reveals the intended answer.
Record recruitment criteria, participant context, task materials, facilitation rules, completion data, group labels, item moves, comments, and accessibility accommodations. AI can summarize patterns after identifiers are removed, but researchers should inspect the raw results and outliers.
Do not collapse all participants into one “average user.” New customers and expert administrators may use different mental models. A disputed card can signal multiple valid paths, a vague label, a composite task, or a recruiting difference. Keep those hypotheses for the next test.
Tree testing removes page design and asks participants where they would go in a text hierarchy. Write scenarios that describe a goal without repeating the target label. Define the correct destination and acceptable alternative paths before running the study.
Measure completion, directness, first choice, backtracking, time where appropriate, and qualitative explanations. Break results down by task and relevant audience rather than relying on one overall score. A high average can hide a critical task that nearly everyone misses.
Test multiple entry points when search engines, deep links, account dashboards, or help links are common. The primary menu is only one access route. Do not claim a hierarchy works merely because participants eventually reach an answer after repeated guessing.
Use realistic content to check whether the proposed tree survives implementation. For each priority task, walk from likely entry points to the outcome. Check page purpose, heading hierarchy, link text, breadcrumbs, search terms, redirects, and the next action.
Review keyboard and screen-reader navigation, but do not reduce accessibility to a menu component. Clear labels, predictable structure, descriptive headings, multiple routes, and maintained page relationships all contribute to orientation. Include people with relevant access needs in research.
Run structural checks for orphan pages, circular paths, identical labels with different meanings, categories with one unexplained child, pages placed under multiple parents without canonical ownership, and content that no longer serves an approved task. Flag these for people; do not let AI silently “fix” URLs or redirects.
The decision record should include the inventory snapshot, task register version, candidate structures, study plans, participant criteria, results, selected hierarchy, label glossary, cross-links, rejected alternatives, unresolved risks, content-owner decisions, and effective date.
Research owns interpretation of user evidence. Content design owns labels and page purpose. Product and service owners approve scope and tradeoffs. Engineering confirms technical feasibility. Accessibility specialists and participants validate relevant barriers. No model substitutes for these responsibilities.
After launch, monitor on-site search refinements, zero-result searches, support contacts, task completion evidence, common backtracking, broken links, orphan content, and locale-specific problems. Treat a material content, audience, product, or navigation change as a reason to retest, not merely to ask AI for a fresh sitemap.
It can propose a reorganization, but URLs reveal the current implementation rather than necessarily showing user needs. Combine a frozen inventory with real task evidence, require citations, and test the resulting structure.
There is no universal number. Use the shallowest structure that expresses meaningful relationships and supports tested paths. Depth is less important than understandable labels, sensible grouping, and successful task completion.
Not automatically. Journeys describe sequences, while navigation categories help people locate something from different starting points. Test whether lifecycle labels match how your audiences actually seek information.
No. Search depends on content structure, titles, metadata, vocabulary, and maintained destinations. People also browse, arrive through deep links, and need orientation after landing on a page.
Preserve the disagreement. It may reveal different audiences, multiple valid paths, unclear cards, or overlapping concepts. Use follow-up research and tree testing instead of forcing an artificial consensus.
AI may help rephrase a documented hypothesis, but it must not turn it into research evidence. Label hypotheses clearly and validate them with representative users before they shape the architecture.
Maintain approved concepts and locale-specific terms, then test labels with speakers in context. Do not rely on literal translation or assume one locale's category boundaries transfer unchanged.
Review it after material audience, service, content, locale, or platform changes and when evidence shows findability problems. Continue lightweight monitoring between formal research rounds.
Disclaimer: This article provides general information, not a substitute for user research, accessibility assessment, legal review, or professional content and product judgment.
Sources checked 6 September 2026.
Sign up to experience all premium features at no cost.
*Available only to new users. Each user is limited to one trial.