Demystifying Workload Management in Splunk OnPrem Environment: Concepts, Best Practices, and Real‑World Insights
By Kamal Dorairaj, Senior Splunk Consultant
As Splunk deployments grow in scale and complexity, effectively managing system resources becomes critical to platform stability and user experience. Workload Management (WLM) is one of Splunk’s most powerful—but commonly misunderstood—features. This blog walks through how it works under the hood, and how to apply best practices in real-world environments.
Why Workload Management Depends on systemd
On Linux systems, workload management relies on cgroups, which are tightly controlled by systemd. To ensure Splunk can properly manage its own cgroup hierarchy, Splunk must be run as a systemd service with delegation enabled.
In your Splunk.service unit file:
| 1 Delegate=true |
Without this configuration, systemd may arbitrarily move Splunk processes between control groups, undermining workload management entirely. Simply put: no systemd delegation, no effective WLM.
Understanding Workload Categories
Splunk classifies all workloads into one of three predefined categories. These categories are not configurable—Splunk decides automatically based on the process type.
1. Ingest Category (The Foundation)
The ingest category is the backbone of Splunk. It includes:
- splunkd (main process)
- Process runner
- KV Store
- App Services
- Introspection and maintenance tasks
If ingestion fails, Splunk becomes unavailable. That’s why ingestion must always be protected.
Best Practice
- Allocate 100% memory to ingest
- Assign higher CPU weight than search
Example:
| 1. [workload_category:ingest] 2. cpu_weight = 50 3. mem_weight = 100 |
This ensures that, during memory pressure, Linux will terminate search processes—not splunkd.
2. Search Category (The Real Workhorse)
All searches—scheduled, ad hoc, and background—start as a splunkd handler and are classified as search workloads.
Search is where resource contention usually originates:
- Search Heads (SH) perform heavy transformation work
- Commands like stats, eval, and rex often execute on the SH
- The SH aggregates partial results from all indexers
- Poorly optimized searches amplify resource pressure
In Splunk Cloud or large environments, thousands of concurrent searches can hit the SH and indexers simultaneously.
3. Misc Category (Everything Else)
If a process is neither ingest nor search, it automatically lands in the misc category.
Typical misc workloads include:
- Scripted inputs
- Modular inputs
Best Practice:
Search Heads and Indexers should not run modular or scripted inputs. Those belong on Heavy Forwarders (HF).
If no misc workloads exist, the category can be effectively disabled:
| 1. [workload_category:misc] 2. cpu_weight = 0 3. mem_weight = 0 |
Protecting Ingestion Without Starving Search
A common misconception is that ingestion requires significant CPU and memory on indexers. In reality:
- Indexers primarily store events
- Parsing and line breaking should already be done upstream (HF)
- The real indexer workload is serving searches
Allocating ~20% CPU to ingestion on indexers is usually sufficient—as long as ingestion has guaranteed access to resources.
A commonly used ratio:
- 65–70% Search
- 20% Ingest
- Remaining buffer for system overhead
Search Prioritization with Workload Pools
Workload Management allows granular search prioritization using workload pools and rules.
For example, separating security and IT monitoring workloads:
| 1 [workload_pool:security_search_pool] 2 cpu_weight = 20 3 mem_weight = 100 4 5 [workload_pool:it_search_pool] 6 cpu_weight = 20 7 mem_weight = 100 |
Rules map searches into these pools:
| 1 [workload_rule:security_search_rule] 2 predicate = app=es 3 workload_pool = security_search_pool 4 5 [workload_rule:it_search_rule] 6 predicate = app=itsi 7 workload_pool = it_search_pool |
High-priority pools either:
- Receive more resources, or
- Restrict the number of concurrent searches
Search Heads vs Indexers: Different Needs, Different Weights
Yes—you can (and should) allocate different CPU and memory weights per role.
- Search Heads typically require more resources for search workloads
- Indexers require balanced search capacity to serve distributed queries
The key rule:
Pool names must match across the deployment, but resource percentages can differ.
What Workload Management Can—and Cannot—Do
One key takeaway from real customer discussions:
WLM does not reduce resource consumption.
It only determines who gets the resources first.
Think of WLM as:
Changing the speed limits on a highway—not reducing the number of cars.
If indexers are overwhelmed, investigate:
- Long-running searches
- Inefficient SPL
- Cache thrashing
- Oversized search bundles
- User behavior (e.g., many concurrent index=* searches)
Ad hoc searches always take priority and can quickly consume system resources if left unchecked.
Final Thoughts
Workload Management is not a silver bullet—but when configured correctly, it provides predictability, protection, and control.
Key principles to remember:
- Always protect ingestion
- Optimize searches before tuning WLM
- Search Heads do most of the heavy lifting
- Use workload pools for prioritization, not punishment
- WLM guides resource usage—it doesn’t fix bad behavior
Used wisely, Workload Management can be the difference between a consistently reliable Splunk platform and a perpetually overloaded one.
Want to get more control over your Splunk resources? Connect with TekStream to discuss workload management, search optimization and keeping your platform running reliably. Talk to a Splunk Expert!
About the Author
Kamal Dorairaj has over 22 years of diverse IT experience with full cycle development of various Applications and Systems. In his first 10 years; he worked as a CRM Consultant at the client locations includes Middle East, India, US; in various domains includes Retail, Telecom, eCommerce. Later he worked as a Lead SRE at Stubhub (was eBay company) for 7 years as a SME in Splunk, GCP. He is a PMP certified, ITIL certified, CompTia Security+ certified and Splunk Architect, Splunk ES Admin, Splunk Developer certified. He also completed Cybersecurity for Managers, Business Analytics from MIT Sloan and MBA in Technology Management. His core interest area has been, ‘Get more business value out of Big data’.