clone-website
Reverse-engineer and clone one or more websites in one shot — extracts assets, CSS, and content section-by-section and proactively dispatches parallel builder agents in worktrees as it goes. Use this whenever the user wants to clone, replicate, rebuild, reverse-engineer, or copy any website. Also tr
By jcodesmore · 1,413 installs
npx skills add jcodesmore/ai-website-cloner-template --skill clone-website
Source repository · Upstream listing
Clone Website
You are about to reverse engineer and rebuild $ARGUMENTS as pixel perfect clones.
When multiple URLs are provided, preserve every pathname as a distinct route and isolate each target's research, screenshots, components, and assets. URLs that differ only by query string or fragment share a pathname, so resolve their route and state behavior explicitly in the output plan. Parallelize page work only after the shared foundation and output plan are fixed so concurrent builders cannot overwrite one another.
This is not a two phase process (inspect then build). You are a foreman walking the job site — as you inspect each section of the page, you write a detailed specification to a file, then hand that file to a specialist builder agent with everything they need. Extraction and construction happen in parallel, but extraction is meticulous and produces auditable artifacts.
Scope Defaults
The target is whatever page $ARGUMENTS resolves to. Clone exactly what's visible at that URL. Unless the user specifies otherwise, use these defaults:
Fidelity level: Pixel perfect — exact match in colors, spacing, typography, animations
In scope: Visual layout and styling, component structure and interactions, responsive design, mock data for demo purposes
Out of scope: Real backend / database, authentication, real time features, SEO optimization, accessibility audit
Customization: None — pure emulation
If the user provides additional instructions (specific fidelity level, customizations, extra context), honor those over the defaults.
Output Isolation and Route Preservation
Treat every target URL as durable project output, not as permission to replace whatever was built previously.
Choose an <app root before extraction. For a single application, <app root is the repository root ( . ). If different origins need separate applications, require the user to provide or approve a prepared Next.js project root for each origin; verify each root builds independently, and never write one origin's output into another root.
Then assign each target:
A collision resistant <site key : a readable origin slug (including a non default port) plus the first 8 lowercase hex characters of SHA 256 over the normalized origin.
A collision resistant <page key : a segment preserving readable pathname slug plus the first 8 lowercase hex characters of SHA 256 over the normalized pathname and any stateful query/fragment; use root <hash for / . Never rely on lossy character replacement alone.
An artifact root: <app root /docs/research/<site key /<page key / .
A screenshot root: <app root /docs/design references/<site key /<page key / .
A component root: <app root /src/components/sites/<site key /<page key / , with genuinely shared same site components under <app root /src/components/sites/<site key /shared/ .
An asset root: <app root /public/sites/<site key /<page key / , with genuinely shared same site assets under <app root /public/sites/<site key /shared/ .
A Next.js route file.
All paths in the remaining phases are relative to that target's <app root . Before writing, verify that every planned route, artifact root, screenshot root, component root, asset root, and downloader filename is unique or is an explicitly approved shared location.
Routing defaults:
For the first single URL clone in an untouched template, the existing scaffold at src/app/page.tsx may be replaced so the clone remains available at / .
For multiple URLs from the same origin, or any later clone added to a project that already contains cloned/user authored pages, preserve the normalized source pathname as its App Router URL (for example, /docs/intro becomes <app root /src/app/docs/intro/page.tsx ). Encode filesystem segment names that would invoke App Router syntax: escape a leading or @ , and literal parentheses or square brackets, with percent encoded folder spellings rather than creating private folders, slots, route groups, or dynamic segments. Verify the built route resolves at the exact normalized URL before completion.
Inspect every existing src/app/ /page.tsx before writing. Never delete or replace a non scaffold route, component tree, research folder, screenshot, or asset namespace unless the user explicitly approves that exact replacement.
If the planned route already exists, stop and ask whether to update that route, choose another route, or skip it.
URLs from different origins may require incompatible fonts, global CSS, layouts, and metadata. Before modifying files, ask whether the user wants separate prepared application roots (recommended) or an intentionally combined multi site app with route scoped styling. Do not create an unapproved monorepo or silently mix global foundations.
Pre Flight
1. Browser automation is required. Check for available browser MCP tools (Chrome MCP, Playwright MCP, Browserbase MCP, Puppeteer MCP, etc.). Use whichever is available — if multiple exist, prefer Chrome MCP. If none are detected, ask the user which browser tool they have and how to connect it. This skill cannot work without browser automation.
2. Parse $ARGUMENTS as one or more URLs. Normalize and validate each URL; if any are invalid, ask the user to correct them before proceeding. For each valid URL, verify it is accessible via your browser MCP tool.
3. Verify the base project builds: npm run build . The Next.js + shadcn/ui + Tailwind v4 scaffold should already be in place. If not, tell the user to set it up first.
4. Inventory existing routes ( src/app/ /page.tsx ), site component namespaces, research artifacts, screenshots, and public assets. Distinguish the untouched template scaffold from existing cloned or user authored work.
5. Write an output plan listing every target URL, <app root , <site key , <page key , destination route, artifact roots, and whether any shared foundation file must change. Resolve collisions across every planned output, same path query/fragment behavior, and multi origin layout decisions with the user before editing.
6. Create only the planned per page/per site directories plus scripts/ if needed. Use unique asset download script names such as scripts/download assets <site key <page key .mjs ; do not overwrite another page's downloader.
7. For multiple pages from one origin, build the shared foundation once, sequentially, before parallel page work. Optionally confirm whether to run page builders in parallel (recommended if resources allow) or sequentially to avoid overload.
Guiding Principles
These are the truths that separate a successful clone from a "close enough" mess. Internalize them — they should inform every decision you make.
1. Completeness Beats Speed
Every builder agent must receive everything it needs to do its job perfectly: screenshot, exact CSS values, downloaded assets with local paths, real text content, component structure. If a builder has to guess anything — a color, a font size, a padding value — you have failed at extraction. Take the extra minute to extract one more property rather than shipping an incomplete brief.
2. Small Tasks, Perfect Results
When an agent gets "build the entire features section," it glosses over details — it approximates spacing, guesses font sizes, and produces something "close enough" but clearly wrong. When it gets a single focused component with exact CSS values, it nails it every time.
Look at each section and judge its complexity. A simple banner with a heading and a button? One agent. A complex section with 3 different card variants, each with unique hover states and internal layouts? One agent per card variant plus one for the section wrapper. When in doubt, make it smaller.
Complexity budget rule: If a builder prompt exceeds ~150 lines of spec content, the section is too complex for one agent. Break it into smaller pieces. This is a mechanical check — don't override it with "but it's all related."
3. Real Content, Real Assets
Extract the actual text, images, videos, and SVGs from the live site. This is a clone, not a mockup. Use element.textContent , download every <img and <video , extract inline <svg elements as React components. Generate content only when it is clearly server generated and unique per session, or when the optional Atlas Cloud fallback below is explicitly approved after the original asset proves unrecoverable.
Layered assets matter. A section that looks like one image is often multiple layers — a background watercolor/gradient, a foreground UI mockup PNG, an overlay icon. Inspect each container's full DOM tree and enumerate ALL <img elements and background images within it, including absolutely positioned overlays. Missing an overlay image makes the clone look empty even if the background is correct.
4. Foundation First
Nothing can be built until the foundation exists: global CSS with the target site's design tokens (colors, fonts, spacing), TypeScript types for the content structures, and global assets (fonts, favicons). This is sequential and non negotiable. Everything after this can be parallel.
5. Extract How It Looks AND How It Behaves
A website is not a screenshot — it's a living thing. Elements move, change, appear, and disappear in response to scrolling, hovering, clicking, resizing, and time. If you only extract the static CSS of each element, your clone will look right in a screenshot but feel dead when someone actually uses it.
For every element, extract its appearance (exact computed CSS via getComputedStyle() ) AND its behavior (what changes, what triggers the change, and how the transition happens). Not "it looks like 16px" — extract the actual computed value. Not "the nav changes on scroll" — document the exact trigger (scroll position, IntersectionObserver threshold, viewport intersection), the before and after states (both sets of CSS values), and the transition (duration, easing, CSS transition vs. JS driven vs. CSS animation timeline ).
Examples of behaviors to watch for — these are illustrative, not exhaustive. The page may do things not on this list, and you must catch those too:
A navbar that shrinks, changes background, or gains a shadow after scrolling past a threshold
Elements that animate into view when they enter the viewport (fade up, slide in, stagger delays)
Sections that snap into place on scroll ( scroll snap type )
Parallax layers that move at different rates than the scroll
Hover states that animate (not just change — the transition duration and easing matter)
Dropdowns, modals, accordions with enter/exit animations
Scroll driven progress indicators or opacity transitions
Auto playing carousels or cycling content
Dark to light (or any theme) transitions between page sections
Tabbed/pill content that cycles — buttons that switch visible card sets with transitions
Scroll driven tab/accordion switching — sidebars where the active item auto changes as content scrolls past (IntersectionObserver, NOT click handlers)
Smooth scroll libraries (Lenis, Locomotive Scroll) — check for .lenis class or scroll container wrappers
6. Identify the Interaction Model Before Building
This is the single most expensive mistake in cloning: building a click based UI when the original is scroll driven, or vice versa. Before writing any builder prompt for an interactive section, you must definitively answer: Is this section driven by clicks, scrolls, hovers, time, or some combination?
How to determine this:
1. Don't click first. Scroll through the section slowly and observe if things change on their own as you scroll.
2. If they do, it's scroll driven. Extract the mechanism: IntersectionObserver , scroll snap , position: sticky , animation timeline , or JS scroll listeners.
3. If nothing changes on scroll, THEN click/hover to test for click/hover driven interactivity.
4. Document the interaction model explicitly in the component sp