MOSTLY AI platform redesign hero

Rebuilding the MOSTLY AI platformFrom stalled to scalable

Overview

MOSTLY AI's platform fused model training and data generation into one rigid flow. Change a setting, get more records, tweak an output, and you retrained from scratch. Enterprise deals stalled on the rigidity. New users gave up before reaching value.

I led the redesign that pulled the two apart. Train once, generate on demand. That one decision reset how the platform scaled, and how the team shipped.

Role

Lead Product Designer

Duration

Apr 2023 – Feb 2024

8 weeks to design, 30 releases to roll out

Collaborators

Product Management

Engineering, UX Research

UX Writing, QA

Tools

Figma

Hotjar

Mixpanel

The problem

One change meant starting over

Training and generation were locked in one job. No model could be reused.

Model training
Data Generation
Coupled job
Synthetic Data

One path for two kinds of users

Creators could train. Consumers just needed data, and hit the wall.

Data creators
Model training
Synthetic data
Data consumers
Model training

The bet

Separate training from generation

Treat the trained model as an asset, not a step. Train it once. Generate from it forever.

Coupled job
Generator
Synthetic data
Model training
Data generation

Every problem traced back to one fused job — splitting it was the decision the rest of the redesign hung on.

What decoupling unlocked

Train once, generate any time

One trained model. Reused on demand, and open to anyone who needs data.

Once
Any time, anyone
Data creators
Data consumers
Model training
Generation
Synthetic data
Generator

A trained model became a reusable asset. Generate again and again, with no setup for the people who just need data.

From nine screens to eight clicks

Before

Nothing moved without manual setup

Classify every table by hand. Tune model settings. Decode what a "job" or a "catalog" even was. Then wait.

01.Pick a source, upload files
Four ways in, two table types to grasp, all before any data moved.
Pick a source, upload files

The original flow, as I found it. Every synthetic dataset started here.

After

Eight clicks from file to synthetic data

The platform does the thinking now. Auto-detected columns, defaults tuned for accuracy, the primary action always one click away.

01.Upload data
One click from the home page. Your Generator is created for you.

Welcome, Alex 👋

Build with data
Train locally
Train your generator locallyUsing our SDK
Import a trained generatorUpload a locally trained model
Trending resources
mostlyai/US Census dataset97% accuracy · Created 2 days ago9423
sarah/Customer transactions96% accuracy · Created 4 days ago125
david/Retail baskets95% accuracy · Created 1 week ago478
mostlyai/Insurance claims98% accuracy · Created 1 week ago3817
elena/Web sessions96% accuracy · Created 2 weeks ago3114
marco/Loan applications97% accuracy · Created 3 weeks ago269
nina/Telco CDRs95% accuracy · Created 1 month ago196
mostlyai/Survey responses98% accuracy · Created 1 month ago12041
tom/IoT sensor stream96% accuracy · Created 2 months ago177
priya/HR headcount97% accuracy · Created 3 months ago61

Designing the experience

Decoupling made connectors a first-class entity

Once just a row and a setup form. Now an object you can open, with status, details, and the generators that depend on it.

01.The original setup
A connector was a row and a modal form. Nothing to open.
The original connector setup screen

The original connector setup, then the entity page that gave connectors parity with generators and datasets.

Built for volume

When datasets multiply, you manage them like a fleet

Reuse made datasets cheap, so lists grew long. Selection, download, and delete moved to bulk.

Synthetic datasets
Owner:Me
1–34 of 34
Name
Status
Visibility
Created
Activity
Customer transactions
Ready
public
2 hours ago
38
Loan applications
Ready
private
5 hours ago
1037
Patient records
Ready
private
yesterday
1766
Web sessions
Ready
private
2 days ago
2495
Retail baskets
In progress
public
4 days ago
00
Insurance claims
Ready
private
1 week ago
38153
Credit scores
Ready
private
2 weeks ago
45182
Churn cohort
Ready
private
3 weeks ago
52211
Fraud signals
Ready
public
1 month ago
59240
Clickstream events
Queued
private
2 months ago
00
Subscription billing
Ready
private
2 hours ago
73298
Support tickets
Ready
private
5 hours ago
80327
Sensor telemetry
Ready
public
yesterday
87356
Ad impressions
Ready
private
2 days ago
4385
Order fulfillment
Ready
private
4 days ago
1134
Payroll ledger
In progress
private
1 week ago
00
Marketing leads
Ready
public
2 weeks ago
2592
Inventory snapshots
Ready
private
3 weeks ago
32121
Ride trips
Ready
private
1 month ago
39150
Energy usage
Ready
private
2 months ago
46179
Hospital admissions
Ready
public
2 hours ago
53208
Genome markers
Ready
private
5 hours ago
60237
Survey responses
Queued
private
yesterday
00
Mobile events
Ready
private
2 days ago
74295
Payment disputes
Ready
public
4 days ago
81324
Warehouse picks
Ready
private
1 week ago
88353
Loyalty points
In progress
private
2 weeks ago
00
Device logs
Ready
private
3 weeks ago
1231
Account signups
Ready
public
1 month ago
1960
Shipment tracking
Ready
private
2 months ago
2689
Return requests
Ready
private
2 hours ago
33118
Trading orders
Ready
private
5 hours ago
40147
KYC profiles
Ready
public
yesterday
47176
Session replays
Ready
private
2 days ago
54205
0 selected

The decoupling made datasets multiply. The interface had to keep up.

One system, searchable

Decoupling turned the platform into one searchable graph

Every object, generators, connectors, datasets, and more, became a first-class entity. So one search could reach them all.

Production data warehouseConnectorSnowflake
Alex IchimCreated Jan 10, 2024
Financial fraud pattern generatorGenerator
James LiuCreated Mar 22, 2024
Retail customer profiles v2Synthetic Dataset
Alex IchimCreated Jan 28, 2024
Customer churn simulatorGenerator
Alex IchimCreated Jan 15, 2024
Transaction history 2023Dataset
James LiuCreated Apr 05, 2024

Search across eight entity types. Try it.

Async, made visible

When work runs in the background, it has to reach you

Decoupling made training and generation asynchronous. Notifications replaced watching a log with being told when it's done.

Welcome, Alex 👋

Build with data
Train locally
Train your generator locallyUsing our SDK
Import a trained generatorUpload a locally trained model
Trending resources
mostlyai/US Census dataset97% accuracy · Created 2 days ago9423
sarah/Customer transactions96% accuracy · Created 4 days ago125
david/Retail baskets95% accuracy · Created 1 week ago478
mostlyai/Insurance claims98% accuracy · Created 1 week ago3817
elena/Web sessions96% accuracy · Created 2 weeks ago3114
marco/Loan applications97% accuracy · Created 3 weeks ago269
nina/Telco CDRs95% accuracy · Created 1 month ago196
mostlyai/Survey responses98% accuracy · Created 1 month ago12041
tom/IoT sensor stream96% accuracy · Created 2 months ago177
priya/HR headcount97% accuracy · Created 3 months ago61

System-wide alerts, so nothing needs babysitting.

Constraint as a tool

Narrowing the screen showed us what mattered

The old product didn't flex at all. Rebuilding it down to 450px forced a call on every element: essential, or clutter.

/Customer transactions
Customer transactionsPrivate
12
Created byAlex Ichim· 2 hours ago
Description
Add a description…
Accuracy97%
StatusReady
Tables1 table
Synthetic datasets12 datasets
customerstabular
Sample size10,00010,000
Accuracy98.2%
Cosine similarity0.999250.99985
Discriminator AUC57.6%51.3%
Distances0.440.46
customers
IDnameDistrictIDCreatedByUserIDBirthdate
mostly72-06a4-4691-93d5-4fe6278eeeefJohnDE-2593812002-07-28
mostlyde-fcf9-4730-abb7-bc109e4a47f4_RARE__RARE_21973-07-12
mostly92-d298-425c-aa89-fb41b2c0d7a3AndrewUK-WV212002-08-02
mostly24-8914-4d2d-9ee8-9f55a1c3e8b2_RARE_DE-3816221978-01-16
mostly95-6dc1-4d00-b2a2-7c4e9d2f1a6bJasonDE-9973512002-07-24
Showing 5 of 10,000 rows
customerstabularCPU Intel Xeon On-Demand: 14 CPUs, 26GB RAM
Fetch training data
3s
Analyze training data
5s
Encode training data
8s
Train AI model
3m 45s
Analyze data for model report
45s
Generate data for model report
1m 22s
Create model report
45s
All steps completed in 6m 53s
Training logsLive
Step · Train AI model
00:00:00Initializing generator · MOSTLY_AI/Large
00:00:02Loading encoded dataset · 10,000 rows · 9 columns
00:00:05Epoch 1/50 — loss 2.413 · val 2.510
00:00:11Epoch 5/50 — loss 1.842 · val 1.905
00:00:23Epoch 12/50 — loss 1.221 · val 1.288
00:00:48Epoch 24/50 — loss 0.844 · val 0.901
00:01:31Epoch 38/50 — loss 0.562 · val 0.618
00:02:40Epoch 47/50 — loss 0.401 · val 0.452
00:03:21Epoch 50/50 — loss 0.377 · val 0.429
00:03:40Checkpoint saved · model converged
00:03:45Training complete · 3m 45s elapsed
1[
2 {
3 "id": "7b4d280d-21b4-4480-803b-952fcfa6f5f4",
4 "name": "players (after manually edited_ USA80 - Cuba10 - null10) - players",
5 "catalogId": "b538405e-86bc-4365-a577-411b6cdbca78",
6 "order": 0,
7 "maxSampleSize": 19000,
8 "generationSize": 19000,
9 "rootTable": null,
10 "distanceFromRoot": null,
11 "contextTable": null,
12 "defaultSmartSelectColumns": [],
13 "processOrder": null,
14 "subjectPriority": null,
15 "rows": 19000,
16 "generationMethod": "SUBJECT",
17 "sequenceCroppingMode": "CUT",
18 "maxSequenceLength": 1000,
19 "maxEpochs": 100,
20 "trainingBatchSize": null,
21 "modelSize": "M",
22 "trainingGoal": "ACCURACY",
23 "tableSource": {
24 "relativeTablePath": "/e7925d04-8136-486c-b5bc-4b6e2129307f/players (after manually edited_ USA80 - Cuba10 - null10) - players",
25 "fileType": "CSV",
26 "basePath": "/data/uploads"
27 }
28 }
29]
Synthetic datasets12
customers · synthetic 1210,000 rows · 2 hours ago
customers · synthetic 1110,000 rows · 2 days ago
customers · synthetic 1010,000 rows · 3 days ago
customers · synthetic 0910,000 rows · 4 days ago
customers · synthetic 0810,000 rows · 5 days ago
customers · synthetic 0710,000 rows · 6 days ago
customers · synthetic 0610,000 rows · 7 days ago
customers · synthetic 0510,000 rows · 8 days ago
customers · synthetic 0410,000 rows · 9 days ago
customers · synthetic 0310,000 rows · 10 days ago
customers · synthetic 0210,000 rows · 11 days ago
customers · synthetic 0110,000 rows · 12 days ago
0 px

Outcome

One decision reset how the platform scaled

8x

The team shipped in thirty small releases instead of one blocked pipeline.

484%

Retained users grew steadily through the year as the redesign rolled out.

3x

Task completion climbed from 18% to 53% as the redesign rolled out.

AVERAGE CSAT
3.2of 5
before redesign
AVERAGE CSAT
4.7of 5
after redesign

Satisfaction, measured before and after on tracking I put in place: 3.2 to 4.7 of 5.

What I learned

Constraints do the deciding

The tight deadline, the single seam, the 320px screen, each one forced a call on what mattered. When you can't keep everything, you find out what's essential. I stopped fighting limits and started using them.

Serving everyone means hiding the right things

One product held experts and first-timers. The answer wasn't two products, it was smart defaults and power tucked away, ready when asked for. Simplicity wasn't fewer features. It was better decisions about what to show.

The biggest decision was never on screen

Separating training from generation shaped every flow, entity, and screen that followed. The best thing I designed here, no user ever saw directly. Structure came first. The surface followed.

The redesign wasn't a set of better screens. It was the seam no one saw.