Optimizing x264 settings and per-title ladders

原始链接: https://streaminglearningcenter.com/articles/optimizing-x264-settings-and-per-title-ladders.html

This Hacker News discussion centers on an article about optimizing x264 encoding settings and per-title ladders. Key takeaways from the commentary include: * **Performance Trade-offs:** One user summarizes the technical findings, noting that while the proposed optimizations achieve a 33% bandwidth reduction, they require 3.2x longer encoding times to gain a marginal (0.7) improvement in VMAF score. * **Industry Evolution:** A veteran commenter contrasts modern streaming optimization with past practices, noting that older "shiny disc" formats (like Blu-ray) allowed for exhaustive, multi-pass, scene-by-scene manual encoding that is no longer feasible in the automated streaming era. * **Technical Critiques:** Users also critiqued the article’s presentation, specifically pointing out poor image quality and CMS issues on the source website. Additionally, participants questioned how these aggressive encoding settings might impact decoding performance on older hardware, with the consensus suggesting that playback compatibility should remain largely unaffected.
相关文章

原文

If you’re producing solely H.264-encoded video, optimizing your encodes can save you bandwidth costs and deliver a higher quality experience to your viewers. If you’re considering adding HEVC or AV1 to the mix, you have another reason.

Before you attempt to compute the bandwidth savings and quality improvements these codecs deliver, you should know just how much quality and bitrate you can get from your current H.264 encoder. Otherwise, some of the gain you credit to the new codec is really the gain you left on the table with H.264.

To illustrate this, I produced two series of tests with x264 and FFmpeg. First was to identify the optimal encoding parameters for x264 using 1080p clips, then to apply those parameters to a full encoding ladder and compute the overall benefits. For you TL/DR aficionados, the results come first. The rest of the article explains how we got them.

You can download a Zip file with much of the results of this testing here. Details of what’s in the Zip are below.

The Results (TL/DR)

You’re going to have a ton of questions, but this is the TL/DR section; we’ll answer most of them below.

Figure 1. What each step adds to average quality, average bitrate, and encoding time, from x264 defaults on the Netflix 2015 ladder to tuned x264 on a per-title ladder. Click the image to view at full resolution.

The TL/DR bottom line? Optimizing x264 boosts VMAF by .66 VMAF points while reducing bandwidth from 4.9 Mbps to 3.27 Mbps, or 33%.  Encoding time is more than tripled. As you can see if you peek ahead to Figure 4, the bandwidth savings exceed encoding costs at 186 hours of viewing, making the decision to optimize a no brainer for all but  the tiniest of encoding shops.

Figure 1 starts with x264 at mostly default settings (with a two-second GOP) using a fixed H.264 ladder, and adds one optimization at a time: the veryslow preset, a ten-second GOP, two-pass encoding, three reference frames, and finally a per-title ladder with the same settings.

Figure 2. The top-heavy distribution pattern.

How did we compute the VMAF and bitrates? It’s complicated. Take a deep breath. Exhale slowly. Repeat.

OK. We computed the numbers assuming the distribution pattern shown in Figure 2, which is called the Top-Heavy distribution pattern in the Streaming Learning Center Bitrate Explorer (SBE) shown in Figure 3. Why didn’t we just average the bitrates and VMAF values for all rungs? For several related reasons.

First, if we average the rung values, we give equal weight to the top rung and the bottom rung. That doesn’t match the reality of most streaming services, particularly in the US, Canada, Europe, and other high-bandwidth regions where the top two or three rungs predominate.

In this case, since we’re comparing a fixed bitrate ladder to an optimized per title ladder, the per-title bottom rungs have much higher quality than the fixed ladder because they’re encoded at higher resolutions with optimized parameters. You see this in Figure 3, where the bottom rung of the optimized ladder (red) has a VMAF score that’s 46 points higher than the fixed ladder, much greater than the modest bitrate differential.

If we averaged the rung values, the bottom rung would significantly boost the overall average, which makes no sense given that only a very tiny percentage of viewers would actually watch that rung. Using the Top-heavy distribution pattern, we weight this rung by a much more appropriate .34%.

Figure 3. Fixed ladder vs. per-title ladder for the test clip Meridian.

While we’re staring at Figure 3, let’s look at the other end of the spectrum, the top rung. Here we see that the top rung of the fixed ladder, encoded at around 5.6 Mbps, has a VMAF score of 96. In contrast, the optimized ladder is encoded at about 1.8 Mbps with a VMAF score of 93.

Which looks better to the viewer? Actually, neither, most viewers will consider them about the same. There’s substantial research showing that once a video achieves a VMAF score of 93, additional quality isn’t discernable by the viewer. One key downside of a fixed ladder is that easy-to-encode files often produce one or more rungs that exceed VMAF 93, which wastes storage and bandwidth. This is why we tuned our optimized per-title ladders to a top rung value of 93.

In comparing the fixed ladder vs per title ladders, using the top-heavy distribution assumes that 71.6 % of viewers watch that top rung, which significantly boosts the VMAF average for that ladder, even though viewers can’t actually perceive the quality difference between 93 and 96.

Just remember a couple of things when you look at the .66 VMAF differential between the baseline and optimal ladders. First, had we averaged the rungs rather than assuming the Top-heavy distribution pattern, the difference would have been much higher (6.5 VMAF points, actually). Second, the .66 delta using this distribution pattern is understated because we’re not accounting for the fact that values over 93 aren’t perceivable by the viewer.

OK, breath again. Math lesson over.

Encoding Cost and Breakeven

As you can see in Figure 1, the optimized ladder increased encoding time by 3.2x. To convert that to dollars, we priced the encoding on an AWS c7i.8xlarge instance at $1.428 per hour, which has the same 32 threads as our test workstation, and bandwidth at $0.02 per GB. Our times were measured on the workstation, so treat the encoding costs as estimates. Scaled from our two-minute clips, the baseline ladder cost about $1.25 to encode per hour of content, and the optimized ladder, at 3.2 times the encoding time, cost about $3.99, an extra $2.74 per hour of content.

The Breakeven tab in SBE enables users to compute the breakeven on encoding decisions like the optimized ladder by factoring in bandwidth savings, distribution cost, and encoding cost, using multiple distribution profiles. Figure 4 compares the baseline and optimized ladder using the Top-heavy distribution.

Not surprisingly, BBE shows the same numbers as Figure 1 for bitrate  (4.90 Mbps to 3.269 Mbps) and VMAF  (90.44 to 91.10). At this bitrate differential, and assuming a $0.02/GB bandwidth cost, that saved about $0.0147 in bandwidth for every hour viewed. At that rate, each hour of encoded content must be watched about 186 times before the bandwidth savings cover the extra encoding cost. After a million hours of viewing bandwidth saved equaled $14,682.

Figure 4. Breakeven between the baseline and optimized ladders in SLC Bitrate Explorer. The encoding cost is recovered in 186 hours of viewing, while overall ladder quality improved by .66 VMAF (at a 33% lower bandwidth).

For an audience with more viewers on slower connections, the quality gain would be larger. If we switched the audience to the mobile preset, the VMAF differential jumps to 8.61. However, because this group retrieved much lower bitrate videos, the bandwidth savings dropped to $0.0023/hour and the breakeven increased to 1,169 hours.

What’s it Mean

As a codec researcher, you hate to spend weeks of testing only to reach the obvious answer; while there are multiple configuration options you can adjust for minor gain, the most effective adjustment you can make is to implement some form of per-title encoding.

Two caveats. First, the benefit from per-title encoding depends upon the efficiency of your current encoding ladder. If the top rung is 5.5 Mbps to 6.5 Mbps or higher for 1080p, you should see substantial bitrate savings. If your current top rung is in the 4.5 Mbps range and below, you may achieve some bandwidth savings, but the primary benefit will be higher quality, particularly in the lower rungs.

Second, clip complexity also matters. If you have a single ladder for all content, from soccer matches to the 6:00 news, a per title ladder should deliver substantial bandwidth savings for news and other easy-to-encode clips and and higher quality at similar or slightly higher bitrates for the hardest content. If you’re encoding primarily sports and other hard-to-encode content, bandwidth savings will be much less.

Now that you know how the story ends, let’s return to the beginning and tell you what we did and why.

How we tested

We tested in two stages. The first stage encoded 13 test clips at 1080p and changed one parameter at a time to find the optimal setting for encoding time and quality.  The second stage used those results to build full encoding ladders. In that stage, we started with a fixed ladder using mostly x264 defaults, changing to the optimal configuration one option at a time, and finishing with a per-title ladder that used all the optimal settings.

The 13 clips cover animation, sports, primetime drama and music, news and education, and screen content. All quality scores are VMAF, measured at 1080p. All encoding times are wall-clock time on a single workstation (an Intel Core i9-14900 with 32 logical cores).

Stage One: Testing 1080 to Identify the Optimal Configuration Settings

In this stage we encoded each clip using the different configuration options to identify the optimal settings. We encoded each clip at a unique bitrate that produced a VMAF score of ~93 using the x264 slow preset and otherwise mostly default FFmpeg configurations. This kept each clip at a realistic top-rung quality level instead of a fixed bitrate that would be too high for some clips and too low for others.

The baseline configuration for the first series of tests was the veryslow preset, two-pass encoding, and a two-second GOP. These tests also used eight encoding threads, so their encoding times are relative to that configuration. As you’ll see, the baseline for the second series of tests used the medium preset. Why the difference?

Because for the first series of tests, I wanted to test configuration options other than preset using the preset I thought was most likely to be used by streamers producing files for moderate to heavy volume viewing. It made no sense to test 16 reference frames if the preset used 3. On the other hand, for the ladder testing, I wanted to start at a frequently used baseline encode, and many producers still use Medium.

Back to round one. From this baseline we changed one parameter at a time and measured VMAF, the lowest-scoring frame in the clip (a proxy for transient quality problems), and encoding time.

We tested preset, GOP size, one-pass versus two-pass encoding, reference frames, B-frames, the maximum bitrate cap, the VBV buffer size, and thread count. These are settings that many encoding shops adjust, and each one trades quality against encoding time.

I didn’t experiment with psychovisual settings like psy-rd, psy-trellis, and adaptive quantization. These settings are designed to make video look better to human viewers, and they often do that by distributing bits in ways that lower VMAF and PSNR scores. So, a test that improves VMAF scores could actually make the video look worse. For this reason, regarding these configurations, I tested at x264’s default settings, figuring that’s the setting most producers will use.

The results were generally consistent across most clips and genres. For the configuration option, the preset was the largest quality lever. The medium preset delivered 99.2% of the veryslow preset’s VMAF score in 31% of the encoding time, and ultrafast delivered 89.9% in 12% of the time. The slower presets buy small quality gains at high cost, which matters most for content that is encoded once and watched many times.

Figure 5. Encoding time, VMAF, and lowest-frame VMAF for each x264 preset, as a percentage of veryslow. Average of 13 clips at 1080p.Charts for the other variables are available in the download.

In other findings:

GOP Size: FFmpeg’s x264 default is a maximum GOP of 250 frames, with extra I-frames inserted at scene changes. Streaming needs fixed-length segments, so we disabled scene-change detection and tested GOPs of one, two, five, and ten seconds, with two seconds as the reference. A ten-second GOP improved VMAF by 0.9 points over a two-second GOP, and a five-second GOP improved it by 0.7 points. A one-second GOP cost 1.5 points. Longer GOPs also took longer to encode, about 17% longer at ten seconds. The baseline ladder used two-second GOPs and the optimized ladders used ten.

Single vs. Two-Pass Encoding: FFmpeg encodes using a single pass by default. We compared one-pass and two-pass encoding using the veryslow preset. On average, two-pass improved VMAF by 0.27 points while increasing encoding time by about 13%. We used two-pass in the optimized ladders, where it improved the quality delivered to viewers on all eleven clips by about half a VMAF point on average, with delivered bitrate about 1% higher.

Reference Frames: The veryslow preset uses 16 reference frames by default (the medium preset uses 3). We tested 1, 2, 3, 4, 5, 8, and 16. Reducing reference frames from 16 to 3 cut encoding time by 31% and cost 0.07 VMAF points. We used 3 reference frames in the optimized ladders, where it cut encoding time from 5.8 to 3.9 the baseline and cost 0.13 VMAF points of delivered quality.

B-Frames: The veryslow preset uses eight B-frames by default. We tested 0, 1, 2, 3, 5, 8, and 16. B-frames made little difference to quality between two and sixteen, all within 0.03 VMAF points of the default. Dropping to two B-frames cut encoding time by about 6% at no measurable quality cost, while disabling B-frames entirely saved 15% but cost 0.5 VMAF points. We didn’t adjust B-frames in the optimized ladders to harvest this 6% because it was measured at sixteen reference frames and may not carry over once references are reduced.

Maxrate: x264 doesn’t cap the bitrate by default, but streaming encodes need a cap, so every test used a maximum bitrate of 200% of the target with a 200% buffer, and we tested caps of 100%, 110%, 150% and 400%. Between 150% and 400%, average VMAF changed very little, but the lowest-scoring frame improved steadily as the cap loosened, from 94.39% of the best result at 150% to 97.16% at 200% and 100% at 400%. The 100% and 110% caps undershoot their target, so we requested 110% of the target bitrate in those tests. They overshot by 7% to 9%, so their slightly higher VMAF scores come from extra bits, not the tighter cap. Concerned about bitrate spikes in the streams, we used the 200% cap in every ladder, including the baseline.

VBV Buffer:x264 doesn’t set a VBV buffer by default. With the cap at 200%, we tested buffers of 100%, 200%, and 400% of the target bitrate. A 100% buffer cost 0.5 VMAF points compared to the 200% buffer, and a 400% buffer gained 0.1 points. We used the 200% buffer in every ladder, including the baseline.

Thread Count: By default, x264 picks its own thread count based on the number of CPU cores. We tested 2, 4, 8, 16, and 32 threads. Thread count affected encoding time with limited quality impact (max .29 VMAF points). Limiting x264 to eight threads slowed encoding by 1.7x and improved VMAF by only 0.18 points. The stage-one tests ran at eight threads, so their encoding times are relative to that configuration. We ran all the ladder tests, baseline and optimized, at the default thread count.

The results also varied by content. Sports clips showed the largest quality degradation from faster presets and by disabling B-frames, and primetime content showed the least. One screen-content clip produced a single corrupted frame at the slow preset, which pulls down the lowest-frame line in Figure 5. It is specific to that clip.

As you’ll read about below, you can download a PDF with a time-versus-quality chart for each parameter, with results broken out by genre, and notes on the outliers.

Stage Two: Comparing Fixed and Optimized Ladders

The second stage tested complete encoding ladders to measure the real world impact of our encoding decisions. Some tough decisions in this section, so buckle up.

Our Fixed Bitrate Ladder

As mentioned, when it comes to comparing per-title and fixed ladder encodes, your starting point, the fixed ladder, has tremendous impact on your findings. Use the notoriously expensive ladder in the Apple HLS Authoring Specifications and your per-title technology looks like a life saver. But you’ve gotta have billions in the bank to use that ladder (like Apple) and I’m guessing even Apple doesn’t use it.

I wanted to use a starting point that was real world and looked it. So I went back to the H.264 ladder that Netflix used before implementing per title and made two adjustments. First, I converted all 4:3 resolutions to 16:9. Then, I dropped the rungs crossed out in Figure 6. The reason for the first adjustment is obvious; the second eliminated rungs probably not required for a high level of QoE and made the encoding time comparison much more relevant.

Figure 6. This is the fixed ladder we used for our baseline testing, after deleting the indicated rungs and converting all 4:3 resolutions to 16:9.

As mentioned above, we used a GOP setting of 2-seconds because no producer uses 250 frames, and the Medium preset. Then we added the settings that earned a place in stage one, one at a time and each on top of the last: the veryslow preset, a ten-second GOP, two-pass encoding, and three reference frames. The reference-frame step is a cost step. It is there to recover encoding time, not to improve quality.

Our Per-Title Ladder

The final step replaces the fixed ladder with a per-title ladder using the same encoder settings. Creating the per-title ladder involves two steps. First, the encoder searches for the 1080p bitrate that produces a VMAF score of 93 for that clip, which becomes the top rung. The search could go as high as 8 Mbps, above the fixed ladder’s 5.8 Mbps top rung, so difficult clips could use the bits they needed to reach VMAF 93. One clip, Jazz, used 6.8 Mbps. Elektra reached the 8 Mbps limit at VMAF 92.4, so its top rung sits just below the target.

Then it builds the ladder downward, reducing the bitrate to 60% of the previous rung at each step. Then it encodes two rungs at that bitrate, one at the resolution of the previous rung, the other at one resolution lower. The rung with the higher VMAF score is added to the ladder and the process rinses and the process repeats until the next rung would fall below 235 kbps. The resulting ladders had between four and seven rungs, with five on most clips.

The per-title search itself takes time, because it encodes and scores multiple candidates that are discarded. We didn’t include the encoding times from discarded files in the results. That’s because few companies will run a per-title search like ours, though Netflix’s original encoding technique was even more extensive and may still be. However, most streamers will use the per-title features in services like Bitmovin or AWS Elemental, which run their own analysis at a cost premium over fixed ladder encoding. For this reason, the encoding times in Figure 1 count only the final ladder encodes.

Finally, the ladder tests used 11 of the 13 clips. We dropped the two screen-content clips because their top rungs are around 300 kbps, which leaves room for only one or two rungs in a per-title ladder.

Other Findings

Quality and bitrate results: The results varied by clip. On six clips, the optimized ladder cut delivered bitrate by 44% to 64% while reducing quality by 0.4 to 1.7 VMAF points, nearly all of it at the top rung. Meridian, shown in Figure 3, is the clearest case. Viewers watching the top rung drop from 96 to 93 VMAF, which they shouldn’t notice. However, as previously mentioned, viewers watching lower rungs of the optimized ladder gain as much as 46 points.

On two clips the optimized settings reduced bitrate and improved quality. On the three hardest clips, Elektra, Jazz, and RiverPlate, the optimized settings used 2% to 5% more bitrate and delivered 1.4 to 6.3 more VMAF points, because the fixed ladder’s 5.8 Mbps top rung limited quality on that content.

You see this with the Electra test clip in Figure 7, where the top rung of the optimized ladder topped out at 7,691 kbps while the fixed ladder stopped at 5.8 Mbps and about 1.7 VMAF points lower. Note that the higher rates for per-title is a configuration decision, not an inherent deficit of per-title. With my inhouse system, as well as Bitmovin’s and AWS’ per-title outputs, you can set the maximum bitrate though obviously this limits quality as well.

Figure 7. On some clips, the per-title ladder exceeded the bitrate of the fixed ladder.

Encoding time: The encoding time in Figure 1 moved in both directions. The tuning steps raised it to 5.8 times the baseline, the reference-frame step reduced it to 3.9 times, and the per-title ladder brought it to 3.2 times, mostly because it has fewer rungs and lower bitrates. In retrospect, the 6% encoding time savings delivered by dropping to two B-frames would have been worth a try.

Veryslow preset: Interestingly, the veryslow preset’s gain was not evenly distributed across the ladder. It improved the lower rungs by about 1.7 to 1.9 VMAF points and the top rung by only 0.7 points. That is mostly a property of VMAF: at a score of 95, there is little room left on the scale, so real efficiency gains at the top rung show up as bitrate savings rather than higher scores.

Measured in bitrate, veryslow needed about 10% to 11% fewer bits than medium to match its quality at every resolution we could check. When you evaluate settings for your top rungs, measure the bitrate needed to reach a target quality, not the score at a fixed bitrate.

Figure 8. One surprising finding was that the veryslow preset delivered more quality improvements at lower resolutions.

Two-pass encoding. Two-pass encoding improved delivered quality on all eleven clips, though by as little as 0.1 VMAF points on some. Gains from every setting varied by content, which is why these tests should be run on your own content before you change a production configuration. The average is a useful guide, but the clips at the edges of the range are the ones that cause problems in production.

How to Use This Info

Many readers have done all this and more with H.264, and to those we say congrats and thanks for reading. If you have any tips you’d care to share, send them my way ([email protected]).

If you haven’t optimized your H.264 encodes using a technique like this, you’re almost certainly leaving bandwidth money on the table. The sooner you get started, the sooner you start saving.

If you’re considering adding another codec to the mix and are trying to identify the bandwidth savings, do this H.264 optimization first. Otherwise, your estimated savings will be overstated, because they include H.264 inefficiencies you haven’t fixed yet. In our tests, tuning x264 and moving to a per-title ladder cut delivered bitrate by a third, savings you’d wrongly credit to the new codec.

Obviously, to be fair to the new codec, you should also optimize its settings before computing the savings. To assist your efforts, my next article on this topic will cover HEVC, then SVT-AV1.

What’s in the PDF

The PDF includes the full results from both stages:

  • 1080p tests: a time-versus-quality chart for each parameter, with results broken out by genre, and notes on the outliers.
  • Ladder tests: Figure 1, results for each clip and genre, the ladder definitions, and the methodology.
  • How we ran the analysis in SLC Bitrate Explorer, with screenshots.

You can download the zip file here. It includes a 29% coupon for SLC Bitrate Explorer, the tool we used to build and compare the ladders in this article. If you want this analysis run for your own service, contact us at [email protected].

联系我们 contact @ memedata.com