How to Build Website Information Architecture with AI

How to Build Website Information Architecture with AI

Olivia Park
September 6, 2026· 11 min read

To build website information architecture with AI, begin with evidence of what people actually need to do, not a blank sitemap prompt. Freeze a content inventory and a register of user tasks, then let AI propose groupings, hierarchy, labels, and alternative paths. Treat every proposal as a hypothesis. Card sorting, tree testing, task walkthroughs, accessibility review, and accountable owners determine what is adopted.

The general AI workflow explains how to constrain evidence and verify output. This guide applies it to information architecture, or IA: the organization, naming, and pathways that help people find and understand website content. It does not allow AI to invent user needs or declare an untested navigation successful.

Key Takeaways

  • Derive tasks from research, search logs, support questions, and observed behavior.
  • Keep tasks, content, pages, features, labels, and navigation components distinct.
  • Ask AI for several candidate structures and cite the evidence behind each choice.
  • Preserve conflicting mental models rather than averaging them away.
  • Test findability with representative participants and realistic tasks.
  • Version the approved IA and retain unresolved evidence and exceptions.

How do you build website information architecture with AI from real tasks?

A useful project produces a decision trail, not merely a diagram. Its core artifacts are a frozen content inventory, an evidence register, normalized task statements, candidate structures, a label glossary, test protocols and results, an exception log, and an approved version of the hierarchy.

Keep these concepts separate:

ConceptExampleWhy it matters
User taskReplace an expired payment methodDescribes an outcome in the user's language
ContentEligibility rules and instructionsInformation needed to complete the task
PageBilling settings help pageOne possible delivery container
FeatureUpdate-card formInteractive capability, not a category
LabelPayments and billingA signpost that must be understood
Navigation pathAccount → Payments → Update cardOne route, not proof that the task is easy

An organization chart is not automatically an IA. Neither is a database schema, a product team's ownership map, or a list of current URLs. Those inputs may explain constraints, but the structure should support the language and relationships users understand.

Ground the architecture in research evidence

Step 1: Freeze goals, scope, and the current inventory

Write a project header with the site or section, audience, supported locales, business and user outcomes, research period, analytics period, content snapshot date, known technical constraints, approvers, and excluded areas. Record whether the work is a diagnosis, partial redesign, migration, or new structure.

Inventory every in-scope URL or content object. Include title, purpose, owner, format, audience, locale, lifecycle status, inbound links, traffic or search evidence when approved, and duplication or quality notes. Do not delete apparently unused content during discovery; mark it for a separate governance decision.

Freeze the inventory before asking AI to reorganize it. Otherwise new pages, redirects, or editorial changes can make candidate structures incomparable. If the inventory is incomplete, label the gap rather than allowing the model to fill it from general web conventions.

Step 2: Build an evidence register of real user tasks

GOV.UK advises teams to start by learning user needs rather than assuming them.[1] Translate each research observation into a traceable task record, but preserve the original evidence beside the normalized wording.

Useful evidence can include moderated interviews, contextual inquiry, usability sessions, on-site search terms, support tickets, call reasons, feedback, failed journeys, accessibility research, and observed workarounds. A user interview synthesis workflow can help organize notes, but only approved research belongs in the register.

For each task, record:

  • a stable task ID and neutral statement;
  • audience and situation;
  • intended outcome and evidence source;
  • frequency, importance, or severity using defined measures;
  • related prerequisites and follow-up tasks;
  • locale or regional differences;
  • uncertainty, contradictions, and research gaps; and
  • research owner and approval status.

Do not turn a stakeholder request such as “promote premium services” into a user need without evidence. Do not treat a search query as a complete task without context. Do not let AI fabricate a persona, frequency, motivation, or pain point.

Step 3: Map the whole problem before drawing a tree

Users often enter a site partway through a larger problem and continue elsewhere. GOV.UK's guidance on mapping a user's whole problem encourages teams to understand the steps before, during, and after their service interaction.[3] Record dependencies and adjacent tasks without claiming the website owns all of them.

A customer journey map can show sequence and handoffs, while IA describes grouping and findability. Do not convert journey stages directly into top-level navigation. “Discover, decide, buy, use” may be a useful analysis frame yet a poor set of labels for someone who wants to replace a lost card.

Also separate task statements from implementation requirements. User stories and acceptance criteria describe product behavior; they do not automatically determine page boundaries or menu levels.

Step 4: Ask AI for candidate groupings, not one answer

Provide the frozen task register, approved inventory fields, constraints, and an output schema. Remove personal information from research excerpts and use task IDs in the prompt. A bounded instruction can say:

Using only the supplied task and content records, propose three alternative
groupings. For every group and label, cite supporting task IDs. Preserve
conflicting patterns and mark unsupported placement as unresolved. Do not add
users, needs, pages, features, research findings, or success claims. Return a
hierarchy, cross-links, label alternatives, and questions for researchers.

Ask for meaningfully different candidates: perhaps task-oriented, object-oriented, and lifecycle-oriented. Require the model to explain tradeoffs and identify items that need multiple entrances. Reject invented evidence, uncited placements, duplicated IDs, missing inventory items, and a hierarchy deeper than the agreed limit.

NIST's generative AI profile emphasizes lifecycle risk management and evaluation.[4] Here, that means recording the input and model configuration, validating the output schema, testing proposals with people, and monitoring the approved structure after launch.

How should you test, approve, and maintain the structure?

Step 5: Review hierarchy and labels with domain owners

Review candidate structures before user testing. Remove internal department names unless research shows users understand and seek them. Check whether siblings are conceptually parallel, whether labels overlap, and whether one category has become a miscellaneous drawer.

GOV.UK's service-scoping guidance frames services around what users are trying to achieve rather than existing organizational boundaries.[2] Apply that principle without pretending every website is a government service: define the task and its boundaries in the user's terms, then record organizational constraints separately.

For each proposed label, keep a small glossary with definition, included and excluded content, tested alternatives, locale-specific wording, and owner. The localization glossary workflow is useful when labels span languages. A literal translation may merge distinctions or create new ambiguity, so approve labels per locale.

Step 6: Run card sorting and preserve disagreement

Card sorting helps investigate how participants group representative items and what they call those groups. Use open sorting to discover patterns, closed sorting to evaluate defined categories, or a hybrid when both questions matter. Choose cards from real tasks and content; avoid internal jargon that reveals the intended answer.

Record recruitment criteria, participant context, task materials, facilitation rules, completion data, group labels, item moves, comments, and accessibility accommodations. AI can summarize patterns after identifiers are removed, but researchers should inspect the raw results and outliers.

Do not collapse all participants into one “average user.” New customers and expert administrators may use different mental models. A disputed card can signal multiple valid paths, a vague label, a composite task, or a recruiting difference. Keep those hypotheses for the next test.

Step 7: Tree-test findability with realistic tasks

Tree testing removes page design and asks participants where they would go in a text hierarchy. Write scenarios that describe a goal without repeating the target label. Define the correct destination and acceptable alternative paths before running the study.

Measure completion, directness, first choice, backtracking, time where appropriate, and qualitative explanations. Break results down by task and relevant audience rather than relying on one overall score. A high average can hide a critical task that nearly everyone misses.

Test multiple entry points when search engines, deep links, account dashboards, or help links are common. The primary menu is only one access route. Do not claim a hierarchy works merely because participants eventually reach an answer after repeated guessing.

Step 8: Walk through content, accessibility, and edge cases

Use realistic content to check whether the proposed tree survives implementation. For each priority task, walk from likely entry points to the outcome. Check page purpose, heading hierarchy, link text, breadcrumbs, search terms, redirects, and the next action.

Review keyboard and screen-reader navigation, but do not reduce accessibility to a menu component. Clear labels, predictable structure, descriptive headings, multiple routes, and maintained page relationships all contribute to orientation. Include people with relevant access needs in research.

Run structural checks for orphan pages, circular paths, identical labels with different meanings, categories with one unexplained child, pages placed under multiple parents without canonical ownership, and content that no longer serves an approved task. Flag these for people; do not let AI silently “fix” URLs or redirects.

Step 9: Approve and version the architecture

The decision record should include the inventory snapshot, task register version, candidate structures, study plans, participant criteria, results, selected hierarchy, label glossary, cross-links, rejected alternatives, unresolved risks, content-owner decisions, and effective date.

Research owns interpretation of user evidence. Content design owns labels and page purpose. Product and service owners approve scope and tradeoffs. Engineering confirms technical feasibility. Accessibility specialists and participants validate relevant barriers. No model substitutes for these responsibilities.

After launch, monitor on-site search refinements, zero-result searches, support contacts, task completion evidence, common backtracking, broken links, orphan content, and locale-specific problems. Treat a material content, audience, product, or navigation change as a reason to retest, not merely to ask AI for a fresh sitemap.

Where do IA projects usually go wrong?

  • Starting with “make me a sitemap”: build a traceable task and content evidence set first.
  • Mirroring departments: group around user understanding and record ownership separately.
  • Turning a journey into navigation: use journeys for sequence and IA for grouping.
  • Accepting one AI structure: compare alternatives and preserve conflicting mental models.
  • Testing labels without tasks: use realistic scenarios and predefined destinations.
  • Optimizing only the main menu: include search, deep links, cross-links, and contextual entry points.
  • Calling eventual success a pass: inspect directness, first choice, backtracking, and explanations.

Summary

  • Freeze goals, evidence periods, and a complete content inventory.
  • Derive task statements from approved research and observed behavior.
  • Ask AI for cited candidate structures, alternatives, and unresolved questions.
  • Review hierarchy and labels without importing the organization chart.
  • Validate through card sorting, tree testing, task walkthroughs, and accessibility research.
  • Approve and monitor a versioned IA with accountable human owners.

Frequently asked questions

Can AI create a sitemap from my current URLs?

It can propose a reorganization, but URLs reveal the current implementation rather than necessarily showing user needs. Combine a frozen inventory with real task evidence, require citations, and test the resulting structure.

How many levels should a website hierarchy have?

There is no universal number. Use the shallowest structure that expresses meaningful relationships and supports tested paths. Depth is less important than understandable labels, sensible grouping, and successful task completion.

Should the main navigation follow the customer journey?

Not automatically. Journeys describe sequences, while navigation categories help people locate something from different starting points. Test whether lifecycle labels match how your audiences actually seek information.

Can search replace information architecture?

No. Search depends on content structure, titles, metadata, vocabulary, and maintained destinations. People also browse, arrive through deep links, and need orientation after landing on a page.

What if card-sorting participants disagree?

Preserve the disagreement. It may reveal different audiences, multiple valid paths, unclear cards, or overlapping concepts. Use follow-up research and tree testing instead of forcing an artificial consensus.

May AI write user tasks from stakeholder ideas?

AI may help rephrase a documented hypothesis, but it must not turn it into research evidence. Label hypotheses clearly and validate them with representative users before they shape the architecture.

How do I validate localized navigation labels?

Maintain approved concepts and locale-specific terms, then test labels with speakers in context. Do not rely on literal translation or assume one locale's category boundaries transfer unchanged.

When should the IA be reviewed again?

Review it after material audience, service, content, locale, or platform changes and when evidence shows findability problems. Continue lightweight monitoring between formal research rounds.

Disclaimer: This article provides general information, not a substitute for user research, accessibility assessment, legal review, or professional content and product judgment.

Sources

  1. GOV.UK Service Manual, Start by learning user needs: https://www.gov.uk/service-manual/user-research/start-by-learning-user-needs
  2. GOV.UK Service Manual, Scoping your service: https://www.gov.uk/service-manual/design/scoping-your-service
  3. GOV.UK Service Manual, Map a user's whole problem: https://www.gov.uk/service-manual/design/map-a-users-whole-problem
  4. National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1): https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence

Sources checked 6 September 2026.

Related articles

Start your 3-day free trial

Sign up to experience all premium features at no cost.

*Available only to new users. Each user is limited to one trial.

How to Build Website Information Architecture with AI | AethoVPN