Melaya Mobile · App Testing

Automate app testing on a real phone.Not an emulator guessing.

Agents walk your real user journeys on real devices: install, onboard, tap through, screenshot, report. Devs verify push notifications, deep links, and cross-device behavior on every build. PMs run A/B variants on physical hardware before release and get side-by-side recordings. It is your actual app on an actual phone, not an emulator guessing.

01
// What breaks today

Manual workflows cost more than the agent does.

Three pains every sales and BD team hits weekly. Each one is what your reps actually complain about, not what a feature page would call them.

  1. 01

    Manual QA on real devices is slow and expensive.

  2. 02

    Bugs in push notifications, deep links, and onboarding flows ship because nobody had time to walk every journey on every phone.

02
// Pipelines you can build

Agent workflows: compose, approve, replay.

Every pipeline below is a shape you wire on the canvas using the crew and tools further down. Not a feature we ship for you, a pattern you configure.

P01

Real user journeys, walked

Install, onboard, tap through, screenshot, report. Agents walk your real user journeys on physical hardware and file replayable evidence.

P02

Push, deep link, cross-device checks

Verify push notifications, deep links, and cross-device behavior on every build, before your users do.

P03

A/B variants on hardware

Run A/B variants on physical devices before release and get side-by-side recordings for the decision.

03
// The multi-agent crew

Device QA crew

Real personas from the tech_team crew. Each ships with a tuned system prompt and a default tool allowlist. Swap models per persona on the canvas.

Tech Lead

TechLead

Plans the journey coverage and reads the evidence per build.

Frontend Engineer

FrontendEngineer

Walks the app UI screen by screen and captures each step.

DevOps Engineer

DevOpsEngineer

Wires runs into your release cadence and files the reports.

UI/UX Designer

UIUXDesigner

Compares variants side by side and flags what regressed.

04
// Scoped tools

Tool allowlists: only the actions you grant.

Every tool below is a real shared tool from the Melaya bundle. Allowlist per agent; HITL-gate the writes; revoke any of them in one click.

shared/tools/phone/

Drives your app on a real device: taps, swipes, screenshots, step logs.

phone_open_appphone_get_screen_treephone_screenshotphone_tapphone_swipephone_batch
shared/tools/project_mgmt/

Files what it finds where your team already works.

jira_create_issuelinear_create_issue
shared/tools/knowledge/

Turns run evidence into a searchable record per build.

build_knowledge_from_textbuild_knowledge_from_file
shared/tools/core/

Reads configs and logs alongside the on-device run.

file_readgrep_search
05
// Three knowledge layers

The crew reads what you give it.

Every pipeline ships with three layers of knowledge access. Mix and match per agent on the canvas. No shared vector space with another tenant, no surprise reads, no opaque retrieval.

L1

Static context

includeContext

Per-pipeline documents appended to specific agents' input on every run. The ICP brief, playbook, pricing sheet, or won-deal email corpus. Whatever needs to be there before the agent thinks. You pick which personas get which docs.

L2

RAG retrieval tool

rag_retrieve

A scoped tool granted per-agent. When the agent decides it needs more depth, it queries the workflow's vector store on demand. Same knowledge base as Static context, accessed only when the model asks for it.

L3

Cross-run memory

pipeline_memory

Pipeline-level state that carries from one run to the next. Yesterday's research is in scope for today's follow-up. The crew remembers what it already prospected, what got approved, what was sent. The audit log is the second-order knowledge base.

07
// FAQ

AI agent questions we get every week.

Is this an emulator?

No. It is your actual app on an actual phone. Agents walk the real UI and file replayable evidence: screenshots, step logs, repro scripts.

Do I need to instrument my app?

No instrumentation and no test-framework integration. The agent drives the app through the screen, the way a user does.

Why use Melaya instead of Appium or Maestro scripts?

Scripted frameworks like Appium and Maestro depend on selectors that break when your UI shifts. Melaya's Device QA crew reads the actual screen: the Frontend Engineer agent walks the app visually, so a moved button or copy change does not kill the run. For pixel-exact assertions in a mature CI suite, scripts still earn their place, and many teams run both.

Can n8n or Zapier test a mobile app on a real device?

No. n8n, Zapier, and Make execute predefined trigger-action steps against APIs, and a mobile app under test usually has no API to call. Melaya's Device Control operates a real Android phone: installs the build, opens allowed apps, reads the screen, taps and types, and logs every step. Those tools still win on connector breadth for linear web automations.

How do I prove which user journeys were tested on each release?

Every Melaya run produces full run traces: timestamped steps, per-screen screenshots, and typed failure reasons. The knowledge bundle turns that evidence into a searchable record per build, and the DevOps Engineer agent files reports through the project_mgmt bundle where your team already works. Deterministic-first evaluation runs rule-based checks before any model-graded ones.

Can I reuse the same test journeys on every build?

Yes. Build journeys once on Melaya's canvas of agents, tools, triggers, and approval gates, then save any run as a reusable pipeline. The Tech Lead agent plans journey coverage, the DevOps Engineer agent wires runs into your release cadence, and the same journeys replay on every build with comparable evidence.

How does Melaya test push notifications and deep links on a real phone?

Melaya's Device Control waits for the notification on a real Android phone, taps the notification, and verifies your app lands on the correct screen, capturing a screenshot at each step. Deep links work the same way: the agent opens the link and confirms the destination screen. The core bundle reads configs and logs alongside the on-device run.

What stops an AI agent from doing something destructive in my app?

Human-in-the-loop approval on every write. The Device QA crew walks read-only journeys freely, but any step that publishes, purchases, or sends pauses until a person approves, and the phone stops for on-device approval before publishing. Tool allowlists scope which apps the phone bundle may open, and full run traces show every tap afterward.

Can I compare A/B variants on physical hardware before release?

Yes. Run each variant as a parallel journey on real hardware, and Melaya's UI/UX Designer agent compares the recordings side by side and flags what regressed. PMs get per-step screenshots and step logs from the phone bundle for both variants, so the release call rests on replayable evidence rather than a hallway demo.

Can test runs stay on my own hardware?

Yes. Melaya runs in the cloud or on your local runner, and the test phone connects to that runner, so unreleased builds and test accounts never leave your machines. Per-step model routing across 23 providers lets sensitive steps run on local models while cheaper cloud models write the report.

Build app founders, mobile devs & pms pipelines on Melaya.

Sandbox tier is free with no card. Join the waitlist and we will email you the moment a slot opens.

← Back to every use case
Join the community
// Cookies
Melaya uses a small set of first-party cookies that are strictly necessary to authenticate you, maintain your session, and protect the platform from abuse. We do not use advertising cookies, cross-site trackers, or third-party analytics by default. The full cookie list is in our Privacy Policy.