AWS Transform Now Supports Block Storage Migration to FSx for ONTAP — Benefits and Pitfalls from a Hands-On EC2 Test
## Introduction
This is Part 2 of the series "The Data Foundation for AWS Modernization." Part 1 set out the lens: migration is not the goal but the entry point, and putting FSx for ONTAP at the data foundation lets you re-choose the compute. This part takes that entry point through AWS Transform, the core of modernization, and verifies hands-on that the data area lands on FSx for ONTAP as part of the migration on AWS. Migration from an on-premises VMware environment will come separately; the idea here is to start modernizing from the tightly coupled EC2 × EBS configuration by decoupling the data store.
Until now, when you migrated servers with AWS Transform (MGN), every disk was placed on Amazon EBS. If the source was not ONTAP and you wanted the data area on Amazon FSx for NetApp ONTAP (FSx for ONTAP below), you needed a "two-step migration" — migrate to EBS first, then move the data — or a third-party block storage migration tool.
With the update of 30 August 2026, **FSx for ONTAP can now be chosen directly as the MGN target storage type**.
* AWS Transform announces general availability of Amazon FSx for NetApp ONTAP support: https://aws.amazon.com/about-aws/whats-new/2026/09/aws-transform-fsx-netapp-ontap-support/
That looked useful, so I tried it hands-on straight away.
The source in this test was **not** an on-premises VMware environment but an EC2 instance already on AWS (Amazon Linux 2023). Using that EC2 as the source, I ran the new path that migrates the data area (EBS) directly to an FSx for ONTAP iSCSI LUN.
Running it for real turned up several pitfalls that reading the documentation alone does not reveal.
Besides sorting out what this new feature can and cannot do, this article records the traps I actually stepped on during the test — errors hidden behind a successful job, and an unexpected capacity spike.
Here are the highlights up front, especially the traps to watch for.
* Only **data volumes** can be targeted. Boot is always Amazon EBS.
* What became GA is "the MGN target storage type". It is a different thing from presenting FSx for ONTAP as a VMware datastore (the Amazon EVS side).
* In MGN's phases, "SNAPSHOT" was a volume Snapshot done as a metadata operation, and "LAUNCH" was the creation of a FlexClone (about 44 seconds for 8 GiB).
* **[Trap 1] Finalize is not just clean-up; it is the step where physical capacity temporarily peaks (about 2x)**.
* **[Trap 2]** Even when the job status becomes `COMPLETED`, **a failed final snapshot (fallback) or the creation of an unbootable target can be hiding behind it**.
### Test Assumptions and Scope
The test was run with the following environment and scope.
(Region: ap-northeast-1 / Source: Amazon EC2 (AL2023) / ONTAP: 9.18.1P3D1)
**[What was run]**
The whole flow: agent installation → replication → test launch → cutover → Finalize → teardown.
**[Out of scope for this article (not verified)]**
* Migration with VMware as the source (no vCenter was prepared, so EC2 stood in as the source)
* FSx for ONTAP as an Amazon EVS datastore (a different feature)
* Durations at production-scale data volumes (the test used a tiny 16 GiB environment, dominated by fixed overhead)
* Windows sources, and access over SMB / NFS
* Post-migration optimization with `lun move start`
## The Shape of the Configuration
_Figure 1: three EC2 instances and where each volume goes: the boot column is EBS in all three rows, and only the data column changes to FSx for ONTAP from staging on (dark theme)_
_Figure 2: the control path, from AWS Transform to the ONTAP management endpoint (dark theme)_
The point is that the boot area and the data area go to different places. **The boot column stays Amazon EBS in all three rows, and only the data column changes to FSx for ONTAP from the staging row on.** The target instance boots from EBS and receives the data over iSCSI. The source in this test was EC2, but the lower two rows of the configuration are the same whether the source is a VMware datastore or on-premises block storage.
For the control path in Figure 2, the official blog says a PrivateLink connection is established automatically, but in practice a Network Load Balancer (NLB) and a VPC endpoint service are created inside your VPC. They incur hourly and LCU charges.
Layer | Normal operation (replicating) | After cutover
---|---|---
Source EC2 | Running. The agent keeps sending blocks | Can be stopped
Replication server EC2 | MGN launches it automatically | Terminated by Finalize
Amazon EBS | Boot staging | Becomes the target's root volume
FSx for ONTAP | Staging FlexVol and LUN | Target FlexVol (FlexClone)
Network Load Balancer | Running | Remains after Finalize (manual deletion needed)
Table 1: each layer after cutover: the one that needs manual deletion is the Network Load Balancer
## The Most Important Distinction: What Became GA Is the MGN Target
Mix this up and you will evaluate the wrong product. It is easily confused with "FSx for ONTAP can be used as a VMware datastore".
Question | Current support
---|---
Can FSx for ONTAP be used as the MGN target storage type? | **Yes (GA with this release)**
Is FSx for ONTAP presented as a VMware datastore? | No. The datastore is a separate feature on the Amazon EVS side
Can the boot disk also go on FSx for ONTAP? | No. Boot is always Amazon EBS
Can EBS and FSx for ONTAP be mixed within one server? | No. All data volumes use the same type
Can you migrate agentlessly? | No. Agent-based replication only
Which protocol do clients connect with? | iSCSI. The data is placed as LUNs inside a FlexVol
Table 2: questions and current support: the agentless path is not supported
The Known limitations in the MGN documentation state `Agent-based replication only`. If you were planning around vCenter's agentless path, the plan has to change. The path depends on whether you leave VMs for EC2, or keep the VMs and move only the storage out.
_Figure 3: the FSx for ONTAP settings in the replication template. Only the SVM ID and the Secret ARN are specified, and the note states that boot goes to EBS_
## Environments This Configuration Suits
Whether a migration fits depends on the shape of the workload, not on the product of the source storage.
* The boot / root area and the data area are separate, and the data area is on block storage
* You want capabilities such as Snapshot and thin provisioning on the data area after migration
* You want to turn "migrate, then move the data" into a single step
* You already run ONTAP and want the same mechanisms (Snapshot / FlexClone / SnapMirror) after migration
Most of these apply whatever the source storage is. Even when the source is ONTAP, however, permissions, quotas and Snapshot policies are not migrated and need to be set up again.
### Environments Where Another Option Fits Better for Now
* Boot and data sit on the same volume
* You cannot install an agent (agentless is not supported)
* The source is Windows and you want to move the data as SMB shares
* File-level migration is enough and block-level replication is not needed (a file transfer such as AWS DataSync is sufficient)
* The source is ONTAP and you want to take the volumes as they are (SnapMirror takes fewer steps)
## A Mini Glossary
Term | Meaning | Where it appears in this article
---|---|---
SVM | A logical server inside the file system | Where MGN creates the FlexVol and LUN
FlexVol | A volume that holds data. LUNs live inside it too | MGN creates one for staging and one for the target
LUN | A unit of block storage, exposed over iSCSI | One source disk becomes one LUN
Snapshot | A point-in-time image inside a volume | MGN's SNAPSHOT phase creates it
FlexClone | A writable clone made from a Snapshot | The target FlexVol is one
Split | Detaching a FlexClone from its parent and materializing it | Finalize runs it
## The Procedure and the Measured Time for Each Step
The entry point is the AWS Transform workspace. MGN runs the migration itself, but it is AWS Transform that handles everything from planning to landing zone and network as a single job.
# | Step | Measured
---|---|---
1 | Initialize MGN | Done from the console. The CLI failed (see below)
4 | Install the agent | Failed once (see below)
5 | Initial full sync (16 GiB / 3 disks) | 226 seconds
6 | Test launch | SNAPSHOT 43 seconds, CONVERSION about 4 minutes 30 seconds
7 | Cutover (correct order) | MGN job 649 seconds; 817 seconds from issuing it to confirming boot
9 | Finalize | Split under 60 seconds; about 33 minutes until clean-up completed
(_n=1 measurements at 16 GiB / 3 disks. Fixed overhead dominates, so they cannot be used as a basis for RTO._)
## Snapshot and FlexClone: What AWS Transform's Phase Names Refer To
I lined up the ONTAP side by timestamp to see which ONTAP operation each AWS Transform phase name corresponds to.
AWS Transform phase | Corresponding ONTAP operation
---|---
SNAPSHOT | Creating a volume Snapshot of the staging FlexVol
LAUNCH | Creating a FlexClone from that Snapshot
The SNAPSHOT phase is not a plain data copy but a metadata operation (44 seconds for 8 GiB). The target volume is a clone with `is_flexclone: true`.
What matters most is that **a FlexClone uses almost no physical capacity when it is created**.
Measured just before Finalize, the target volume's logical size was 7.91 GiB while its physical consumption was only 35.5 MiB. Capacity planning has to look at physical values, not logical ones.
## Finalize Is Not Clean-up; It Is Where Capacity Peaks
Running Finalize starts a split of the FlexClone (the operation that detaches it and makes it independent), which materializes the data.
Elapsed from T0 | Observed change
---|---
Immediately | Lifecycle `CUTOVER`, replication `DISCONNECTED`
About 3 minutes | FlexClone split starts. Physical consumption rising from 35.5 MiB → 3.94 GiB
About 4 minutes | Split completes. Physical consumption 8.57 GiB
About 13 minutes | Staging FlexVol deleted, replication server EC2 terminated
There was a gap of about 9 minutes between the split completing and the staging volume being deleted.
**During those 9 minutes, physical capacity holds twice the migrated data.**
If you run Finalize while the aggregate's free space is below the migrated data size, it is likely to get stuck here. Keep in mind that Finalize is a capacity risk, not an availability risk.
## Four Places I Stumbled
### 1. A Failed Final Snapshot Behind a Successful Job
To measure downtime, I shut down the source OS and then ran cutover. The job returned `COMPLETED` and the target `LAUNCHED`. It looked like a success, but the job log recorded this:
09:14:21 SNAPSHOT_START
09:19:22 SNAPSHOT_FAIL timed out after 300 seconds
09:19:23 USING_PREVIOUS_SNAPSHOT
The final sync did not complete, and the job fell back to the previous snapshot. The cause was my procedure: shutting down the OS also stopped the replication agent, and taking the crash-consistent snapshot timed out.
The correct procedure is to stop only the application's writes and run cutover with the OS and the agent still running. "Stop writes on the source" in the official documentation does not mean shutting down the OS.
In this test there were no recent writes, so the data matched, but doing this in a real environment **loses the latest data written after the snapshot it fell back to.** What makes this risky is that it does not show up in the job status (`status` or `launchStatus`) at all. By the API's design, the job returns `COMPLETED` as long as processing finishes, even after a fallback.
**Lesson** : After cutover, always check `event` in `describe-job-log-items`. A `COMPLETED` job is not evidence that the final sync succeeded.
### 2. Two Unbootable Targets: NVMe Device Names Re-enumerated
I ran cutover in the correct order twice, and both produced instances that would not boot (a UEFI shell reboot loop). The job log showed no abnormal events, and the status was `COMPLETED`.
Examining the boot EBS volume of an unbootable instance, **the 8 GiB boot volume held the contents of a 4 GiB data disk.** The boot and data disk assignment had been swapped.
The cause: when the source was stopped and restarted as part of the test, **NVMe device names were re-enumerated and the boot disk moved from`/dev/nvme0n1` to `/dev/nvme1n1`**. MGN's staging-type assignment appears to be keyed on the device name, and it is not re-evaluated after the name moves.
The official best practices say not to restart the source before cutover in the first place. If a restart does happen, check the disk assignment with a command like the one below and redo the test launch.
# Check which device is treated as the boot disk, and its assignment
aws mgn get-replication-configuration --source-server-id "$SRC" \
--query 'replicatedDisks[?isBootDisk==`true`].[deviceName,stagingDiskType]' --output table
### 3. No Repair API, and Three Contradicting Errors
I called `UpdateReplicationConfiguration` to fix the disk assignment, but whatever parameters I passed, it returned contradicting errors, and the API could not fix it.
In the end I deleted the source server and reinstalled the agent, which turned out to be unnecessary: rerunning the installer is the documented procedure.
Rerunning the installer re-establishes the list of replicated disks and their mapping against the actual machine. To avoid losing launch template customizations, try rerunning the installer before deleting anything.
### 4. The Installer's Hint Did Not Match the Real Cause
When the agent installation failed, the installer displayed `Are kernel linux headers installed correctly?`. Reading the log closely, the real cause was lack of space in `/tmp`.
On Amazon Linux 2023, `/tmp` is tmpfs; on a t3.small it has only 955 MiB, which does not meet the installer's requirement (1 GB or more free).
sudo env TMPDIR=/var/tmp/mgn-build TEMP=/var/tmp/mgn-build TMP=/var/tmp/mgn-build \
./aws-replication-installer-init --region ap-northeast-1 --no-prompt
Pointing it at `/var/tmp` like this works around it. The installer's hint is not necessarily the real cause, so check the error code in the log.
## Teardown Checklist
Finalize does not finish the clean-up. The correct teardown order for leaving nothing behind is as follows (steps 1–7 depend on each other).
1. Terminate the target and source EC2 instances
2. Delete the MGN source server
3. Reject the VPC endpoint connection, then delete the endpoint service
4. **Delete the NLB and its target group (important: charges continue until you delete them by hand)**
5. Delete the target FlexVol with `fsx delete-volume`
6. Delete the SVM with `fsx delete-storage-virtual-machine`
7. Delete the ONTAP `security login`
The staging EBS volumes are deleted automatically with a delay, but the target's root EBS volume has no `DeleteOnTermination` and must be deleted manually.
## Initializing AWS Transform from the CLI: Create the IAM Roles First
If running `aws mgn initialize-service` from the API or CLI fails, the cause is that the IAM roles have not been created.
Initializing from the console creates the IAM roles for you, but on the CLI path you need to create eight roles (including `AWSApplicationMigrationFsxProxyRole` for FSx for ONTAP) and attach their policies first.
## Closing
With this GA, a migration now finishes with the data area already on FSx for ONTAP as iSCSI LUNs, and the step of moving the data afterwards is gone.
Any server whose boot and data are separate fits the same shape, whatever the original storage.
On the other hand, watch for the configuration spanning two kinds of storage, the physical capacity spike during Finalize, and MGN-specific traps such as "a job status of COMPLETED does not mean the final sync succeeded".
I hope this test record helps anyone migrating servers with separate boot and data to EC2 while considering FSx for ONTAP as the home for the data area.
## Resources
* AWS What's New: AWS Transform announces general availability of Amazon FSx for NetApp ONTAP support
* MGN: FSx for ONTAP configuration
* MGN: Best practices for AWS Transform MGN
* FSx for ONTAP: Using Amazon Elastic VMware Service with FSx for ONTAP
* Verification report (GitHub): the test environment configuration, the measured teardown procedure, and the list of feedback sent to AWS, all omitted from this article
All test environments have been deleted. The figures are single measurements under a specific environment and conditions, and vary with data size and configuration. The configuration details are in the verification report above.
_This article is Part 2 of the series "The Data Foundation for AWS Modernization." The Japanese original is on hatenablog._