Functional-Testing-for-AI-Native-Workspaces

10+

Core modules validated end-to-end

3

Sensitive data types tested

5

Execution speed modes validated

Overview

mercury.build(Oasis) is a no-code workspace where teams coordinate human and AI agent work together — spinning up agent teams that mix models from Claude, OpenAI, and other frameworks, complete with policies, approvals, and real-time collaboration. Because the platform lets AI agents act semi-autonomously — sending emails, pulling data from connected third-party apps, and executing multi-step workflows — testing it required more than standard functional coverage: every agent action needed to be validated against the permissions, guardrails, and approval logic meant to keep it safe.

Testvox ran an end-to-end functional testing engagement spanning authentication, real-time collaboration, AI agent management, workflow automation, security guardrails, and cross-platform coverage — validating not just that features worked, but that the platform’s safety mechanisms held up under real usage patterns.

Challenges Faced by Client

  • Agents that act on a user's behalf: Oasis's AI agents can retrieve data from connected third-party applications — pulling candidate forms or subscription details from email, for instance — without the user re-authenticating each time. Every one of these flows needed validation to confirm agents only ever retrieved information they were actually authorized to access.
  • Auto-approval features carry real risk if untested: With auto-approval workflows capable of sending emails without human review once a template is trusted, testing needed to confirm the approval logic — manual approval, rejection, and auto-approval — behaved exactly as configured, with no silent gaps.
  • Sensitive data moving through an AI layer: With guardrails designed to detect and mask sensitive information (SSNs, Aadhaar numbers, PAN numbers) before it could be stored or exposed, validating this wasn't optional — it was the core trust boundary of the product.
  • A platform still catching up to itself across surfaces: The browser app was feature-complete, but the desktop app was still under active development and the mobile app wasn't yet available for testing — and PII protection itself was only partially implemented at the time. Testing had to be scoped honestly around what each platform could actually support, rather than applying one blanket test plan everywhere.

Testvox Solution

  1. Full-Lifecycle Functional Coverage
  2. Testvox validated the complete user journey — sign-up, sign-in, account management, room/conversation creation and management, chat messaging (including code blocks, hyperlinks, file attachments), and collaboration panels covering tasks, artifacts, and documents with live real-time sync.

  3. Deep Validation of AI Agent Behavior
  4. Beyond standard feature testing, Testvox validated the full agent lifecycle: creating custom agents from prompts, using prebuilt templates (like email- and invoice-generation agents), importing externally built agents, configuring different AI models and providers, and invoking specific agents by mention inside a conversation — confirming each responded correctly based on its configured permissions and data access.

  5. Guardrail, Permission & Approval Testing as a First-Class Concern
  6. Testvox treated the platform's trust mechanisms as core functional requirements, not an afterthought — validating PII masking across multiple sensitive data types, manual vs. auto-approval workflows for outbound emails, and role-based access control across workspace-, room-, group-, and individual-level permissions, including permissions scoped specifically to AI agents.

  7. Structured Edge-Case & Cross-Platform Testing
  8. Testvox pushed the platform with large payloads, special characters, emoji and multilingual content, and embedded code snippets, while validating response accuracy and timing across all five execution speed modes (Fast, Medium, Slow, High, Maximum) — and clearly scoped platform coverage to reflect what was actually testable across browser, desktop, and mobile at the time.

Result

Agent Permissions Verified at the Data Level

Testvox confirmed that AI agents retrieved only the information they were explicitly authorized to access from connected third-party applications — validating the core trust boundary that lets Oasis's "agents acting on your behalf" model work safely.

Approval Workflows Validated End-to-End

Manual approval, rejection, and auto-approval paths were all confirmed to behave as configured — ensuring outbound actions like automated emails only went out when the platform's own rules said they should.

Clear, Honest Coverage Across a Fast-Moving Platform

Rather than forcing uniform testing onto features still in development, Testvox delivered full coverage where the platform was ready (browser, core workflows) and clearly documented partial coverage where it wasn't (desktop, mobile, PII protection) — giving Oasis's team an accurate picture of readiness, not a false sense of completeness.

Conclusion

Testing an AI-native workspace means testing more than features — it means testing whether the platform's agents, permissions, and guardrails actually hold the line they're designed to hold. Testvox's engagement gave Oasis validated confidence in its core collaboration and agent workflows, precise visibility into its data-protection mechanisms, and an honest map of what still needs coverage as the platform matures.

Related Resources