mirror of https://github.com/kemko/nomad.git synced 2026-01-01 16:05:42 +03:00

Go to file

Tim Gross 0ba7d0036b CSI: persist previous mounts on client to restore during restart (#17840 )

When claiming a CSI volume, we need to ensure the CSI node plugin is running
before we send any CSI RPCs. This extends even to the controller publish RPC
because it requires the storage provider's "external node ID" for the
client. This primarily impacts client restarts but also is a problem if the node
plugin exits (and fingerprints) while the allocation that needs a CSI volume
claim is being placed.

Unfortunately there's no mapping of volume to plugin ID available in the
jobspec, so we don't have enough information to wait on plugins until we either
get the volume from the server or retrieve the plugin ID from data we've
persisted on the client.

If we always require getting the volume from the server before making the claim,
a client restart for disconnected clients will cause all the allocations that
need CSI volumes to fail. Even while connected, checking in with the server to
verify the volume's plugin before trying to make a claim RPC is inherently racy,
so we'll leave that case as-is and it will fail the claim if the node plugin
needed to support a newly-placed allocation is flapping such that the node
fingerprint is changing.

This changeset persists a minimum subset of data about the volume and its plugin
in the client state DB, and retrieves that data during the CSI hook's prerun to
avoid re-claiming and remounting the volume unnecessarily.

This changeset also updates the RPC handler to use the external node ID from the
claim whenever it is available.

Fixes: #13028

2023-07-10 13:20:15 -04:00

.changelog

CSI: persist previous mounts on client to restore during restart (#17840 )

2023-07-10 13:20:15 -04:00

.github

ci: more self-hosted iops for checks workflow (#17852 )

2023-07-10 10:21:04 -05:00

.release

Prepare for next release

2023-06-28 11:06:28 -04:00

.semgrep

[COMPLIANCE] Add Copyright and License Headers

2023-04-10 15:36:59 +00:00

.tours

Make number of scheduler workers reloadable (#11593 )

2022-01-06 11:56:13 -05:00

acl

node pools: list nodes in pool (#17413 )

2023-06-06 10:43:43 -04:00

api

build(deps): bump github.com/hashicorp/cronexpr in /api (#17787 )

2023-07-10 11:23:00 +01:00

[COMPLIANCE] Add Copyright and License Headers

2023-04-10 15:36:59 +00:00

client

CSI: persist previous mounts on client to restore during restart (#17840 )

2023-07-10 13:20:15 -04:00

command

Generate files for 1.6.0-beta.1 release

2023-06-28 11:06:20 -04:00

contributing

Update checklist-rpc-endpoint.md (#17698 )

2023-06-27 10:52:38 +02:00

demo

compliance: add headers with fixed copywrite tool (#17353 )

2023-05-30 09:20:32 -05:00

dev

repo: block pushing to release branches in git hook (#17377 )

2023-06-01 09:36:20 -05:00

drivers

Include parent job ID as a Docker container label (#17843 )

2023-07-10 11:27:45 -04:00

e2e

e2e: respect timeout value when waiting for allocs in v3. (#17800 )

2023-07-10 09:47:10 +01:00

helper

CSI: persist previous mounts on client to restore during restart (#17840 )

2023-07-10 13:20:15 -04:00

integrations

Update metric names (#16894 )

2023-04-18 13:25:42 -07:00

internal/testing/apitests

[COMPLIANCE] Add Copyright and License Headers

2023-04-10 15:36:59 +00:00

jobspec

Add disable_file parameter to job's vault stanza (#13343 )

2023-06-23 15:15:04 -04:00

jobspec2

Add disable_file parameter to job's vault stanza (#13343 )

2023-06-23 15:15:04 -04:00

lib

dep: update from jwt/v4 to jwt/v5 (#17062 )

2023-05-03 11:17:38 -07:00

nomad

CSI: persist previous mounts on client to restore during restart (#17840 )

2023-07-10 13:20:15 -04:00

plugins

Include parent job ID as a Docker container label (#17843 )

2023-07-10 11:27:45 -04:00

scheduler

core: remove unnecessary call to SetNodes and adds DC downgrade test (#17655 )

2023-06-22 13:26:14 -04:00

scripts

[COMPLIANCE] Add Copyright and License Headers (#17732 )

2023-06-26 11:11:17 -05:00

terraform

ci: run 'make check' as reusable workflow (#17600 )

2023-06-20 08:17:13 +01:00

testutil

tests: enable newer windows (#17401 )

2023-06-02 11:38:38 -05:00

tools

tools: update dependencies and use tree set (#16974 )

2023-04-25 07:47:19 -05:00

Report shows a 3rd party browser extension puts a banner at the top of page and awkwardly shifts nav; this fixes that (#17783 )

2023-06-30 17:09:42 -04:00

version

Prepare for next release

2023-06-28 11:06:28 -04:00

website

docs: detail Consul ACL token env var config option. (#17859 )

2023-07-10 14:26:18 +01:00

.copywrite.hcl

build: add agent bindata file to copywrite ignore list. (#17507 )

2023-06-14 11:13:59 +01:00

.git-blame-ignore-revs

add copywrite headers commit to ignore-revs config file (#17037 )

2023-05-01 10:57:43 -04:00

.gitattributes

Remove invalid gitattributes

2018-02-14 14:47:43 -08:00

.gitignore

git: ignore .fleet directory (#16144 )

2023-02-13 07:39:30 -06:00

.go-version

build: update to go1.20.5 (#17451 )

2023-06-07 11:44:59 -04:00

.golangci.yml

[COMPLIANCE] Add Copyright and License Headers

2023-04-10 15:36:59 +00:00

.semgrepignore

build: disable semgrep on structs.go for now

2022-02-01 10:09:49 -06:00

build_linux_arm.go

[COMPLIANCE] Add Copyright and License Headers

2023-04-10 15:36:59 +00:00

CHANGELOG-unsupported.md

docs: split out unsupported versions in changelog (#17704 )

2023-06-23 15:17:57 -04:00

CHANGELOG.md

Prepare release 1.6.0-beta.1

2023-06-28 11:06:05 -04:00

CODEOWNERS

build: update deprecated GitHub Actions (#17218 )

2023-05-17 08:57:28 -04:00

Dockerfile

build: add Docker image (#17017 )

2023-06-23 15:57:09 -04:00

GNUmakefile

ci: remove circleci (#17502 )

2023-06-12 16:28:19 -05:00

go.mod

deps: update cronexpr to capture license file in SBOM tools (#17733 )

2023-06-27 07:58:20 -05:00

go.sum

deps: update cronexpr to capture license file in SBOM tools (#17733 )

2023-06-27 07:58:20 -05:00

LICENSE

[COMPLIANCE] Update MPL 2.0 LICENSE (#14884 )

2022-10-13 08:43:12 -04:00

main_test.go

[COMPLIANCE] Add Copyright and License Headers

2023-04-10 15:36:59 +00:00

main.go

[COMPLIANCE] Add Copyright and License Headers

2023-04-10 15:36:59 +00:00

README.md

Adds public roadmap project to readme

2023-03-20 15:11:38 -07:00

Vagrantfile

dev: make cni, consul, dev, docker, and vault scripts Lima compat. (#16689 )

2023-03-28 16:21:14 +01:00

README.md

Nomad

Nomad is a simple and flexible workload orchestrator to deploy and manage containers (docker, podman), non-containerized applications (executable, Java), and virtual machines (qemu) across on-prem and clouds at scale.

Nomad is supported on Linux, Windows, and macOS. A commercial version of Nomad, Nomad Enterprise, is also available.

Website: https://nomadproject.io
Tutorials: HashiCorp Learn
Forum: Discuss

Nomad provides several key features:

Deploy Containers and Legacy Applications: Nomad’s flexibility as an orchestrator enables an organization to run containers, legacy, and batch applications together on the same infrastructure. Nomad brings core orchestration benefits to legacy applications without needing to containerize via pluggable task drivers.
Simple & Reliable: Nomad runs as a single binary and is entirely self contained - combining resource management and scheduling into a single system. Nomad does not require any external services for storage or coordination. Nomad automatically handles application, node, and driver failures. Nomad is distributed and resilient, using leader election and state replication to provide high availability in the event of failures.
Device Plugins & GPU Support: Nomad offers built-in support for GPU workloads such as machine learning (ML) and artificial intelligence (AI). Nomad uses device plugins to automatically detect and utilize resources from hardware devices such as GPU, FPGAs, and TPUs.
Federation for Multi-Region, Multi-Cloud: Nomad was designed to support infrastructure at a global scale. Nomad supports federation out-of-the-box and can deploy applications across multiple regions and clouds.
Proven Scalability: Nomad is optimistically concurrent, which increases throughput and reduces latency for workloads. Nomad has been proven to scale to clusters of 10K+ nodes in real-world production environments.
HashiCorp Ecosystem: Nomad integrates seamlessly with Terraform, Consul, Vault for provisioning, service discovery, and secrets management.

Quick Start

Testing

See Learn: Getting Started for instructions on setting up a local Nomad cluster for non-production use.

Optionally, find Terraform manifests for bringing up a development Nomad cluster on a public cloud in the terraform directory.

Production

See Learn: Nomad Reference Architecture for recommended practices and a reference architecture for production deployments.

Documentation

Full, comprehensive documentation is available on the Nomad website: https://www.nomadproject.io/docs

Guides are available on HashiCorp Learn.

Roadmap

A timeline of major features expected for the next release or two can be found in the Public Roadmap.

This roadmap is a best guess at any given point, and both release dates and projects in each release are subject to change. Do not take any of these items as commitments, especially ones later than one major release away.

Contributing

See the contributing directory for more developer documentation.

Languages

Go 76.9%

MDX 11%

JavaScript 8.2%

Handlebars 1.7%

HCL 1.4%

Other 0.7%

README.md Unescape Escape

Nomad

Quick Start

Testing

Production

Documentation

Roadmap

Contributing

README.md