Operating parameters for the virtual infrastructure at SCC

A virtualised infrastructure is by nature a »shared-use« system. Many virtual machines (hereafter abbreviated to VMs) share a physical environment, so it is necessary to comply with various parameters and rules for trouble-free operation.

The document is regularly revised and adapted to the respective state of the art. In addition, the most important rules are summarised in this document. These supplement the Operating concept. Every person responsible for a server (also operator or responsible admin) of a VM undertakes to strictly adhere to the points set out here.

The document is regularly revised and adapted to the respective state of the art. In addition, the IT Security Guidelines and the Regulations for the use of the IT infrastructure of the Bauhaus-Universität Weimar apply.

Lifecycle: application, modification, deletion

The central tool for managing the server systems is the SCC server and inventory database (CMDB)
Important information on systems (co-)operated by SCC is stored there. It is the data basis for a number of other processes. The CMDB can be accessed from the Bauhaus-Universität network at:

URL: https://cmdb.scc.uni-weimar.de/
Username: <uni-account>    Password: PW of the Uni account

(Virtual) servers and other systems are managed via this website (and only there). The forms for applying for new systems, changing systems and finally deleting them are electronically mapped there and integrated with a whole series of workflows, e.g. network application for IP/DNS allocation, link to central monitoring and configuration of the backup. Thus, for example, no additional network application for new servers is necessary, as this is done in the background. It is important that changes to servers in operation are registered/applied for in advance via the database interface, as additional work steps are usually required by the SCC to implement the desired adjustments. In justified exceptional cases, this is also possible afterwards, but usually requires additional manual effort. Therefore the request: Always register before making a change! Through the stored workflows, the person responsible for the server also automatically receives feedback on the processing status of the requests made and notices if disruptions or deviations occur in the operation of the VM. The database interface provides the server managers with a summary of the systems under management, notes and overviews of various parameters and system states, as well as consistency and plausibility checks. For VMs, there is also the option of viewing and deleting a list of existing snapshots of the managed systems, see also paragraph »Usage of Snapshots«.

Server administrators who are using the database interface for the first time must be authorised beforehand by entering them in a whitelist and, if necessary, creating firewall rules. To do this, the server administrator must contact vmware[at]scc.uni-weimar.de by e-mail and provide the following data:

  • User account of the Bauhaus-Universität (please do not send a password!)
  • first & last name
  • Divisional affiliation (faculty, department, etc.)
  • Phone number on official business

As soon as an application procedure has been completed, the applicant is automatically informed. For new VMs, a specific user / password combination is delivered depending on the operating system. This initial password must be changed immediately by the person responsible for the server.

Form fields, attributes, classification of the VM into classes

Both before applying for new VMs and during operation, the responsible server operator must think about the correct grouping of his VM into the classes mentioned below. The respective grouping has effects in various areas. If a change is necessary or useful, it must be registered immediately by the server operator via SCC's server database, since such a change almost always requires work processes by SCC. Most attributes should be self-explanatory, more specific things are explained below.

Operator / Deputy (=Responsible 1+2)

These fields designate two persons who are responsible for the corresponding system. This can be of an organisational nature (specialist process manager, technical manager), but a »real« system administrator is preferred. If a system is maintained by external service providers, it is nevertheless essential to specify a university employee here in order to have a contact person on site if necessary.

Operating Systems

The SCC provides a number of pre-installed operating systems as »Base images«. The installation of any other operating system not on the list is not supported by the SCC and is the full responsibility of the server owner(s). If this is absolutely necessary, please select »vendor appliance« as the operating system and contact the virtual infrastructure operator team as indicated in the Contact paragraph. The change of the operating system (change request) may require a complete reinstallation with corresponding data loss, unless an in-place upgrade is possible. The time required for this should be taken into account for any plans.

Deployment / Operating Status

Values: productive, test
Here it is important to distinguish between systems that are used productively (active operation, continuous operation) and systems that are intended as a test for new things or only for short-term use < 90 days.
This classification is important for e.g. backup mechanisms, priorities of resource provision, procedures in case of malfunctions. A test system is often declared productive after a certain time. Systems declared as  »test« are mainly operated at the Coudraystraße location, the »productive« systems mainly at the Steubenstraße location. However, this can deviate in favour of a more balanced load distribution.

Priority

Values: red, yellow, green
This classification is an indicator of the priority of the system and mainly affects processes in the event of malfunctions and failures as well as the allocation of resources within the virtual infrastructure. Systems classified as »test« are usually always »green«. Otherwise, the following table helps with grouping.

PRIOgreenyellowred
Definition for the
allocation
All systems that cannot be classified as yellow or red. Recovery has the lowest priority. In the event of power failures or climatic disturbances, these systems are shut down first. Test systems are usually always green.Permanent (productive) operation of IT services. A failure of these systems has little or no impact on the external image of the university. Internal university processes are massively disrupted. Recovery takes place according to class red systems. In the event of power failures or climatic disruptions, these systems are shut down before those of class red.Permanent (productive) operation of IT services. A failure of these systems severely affects the external image of the entire university. Internal university processes are massively disrupted. Recovery has the highest priority. In the event of power failures or climatic disruptions, these systems are shut down last.
VM Swingneinyes, if planned downtime > 8 hyes, if planned downtime > 4 h
VM Disk
Mirroring
test: no, productive: noproductive: on request, not preferredproductive: on request, preferred
VM CPU sharelowmediumhigh
VM RAM sharelowmediumhigh
BackupAll VMs up to 1.2 TB disk size are backed up daily as an image by default, plus file-based backup with client is possible on request. Physical servers always file-based with client.
Shutdown temperature» 30 °C» 35 °C» 38 °C
UPS residual runtime« 25 Min« 15 Min« 8 Min

 IT Proceedings

Values: https://cmdb.scc.uni-weimar.de/inventar/itverfahren.php
The assignment to an IT procedure serves the organisational grouping of services to the procedures and processes used within the Bauhaus-Universität. If no procedure applies, please contact the staff of the SCC.

Network class

Values: https://cmdb.scc.uni-weimar.de/inventar/netzklassen.php
The assignment to network classes serves the organisational grouping and classification of IT services in the context of IT security and data protection. For example, it is possible for the SCC to operate data protection-relevant services in separate network segments or to assign device classes to special network segments.

Backup

Values: yes (standard), no
Virtual machines are backed up by default. Automatic data backup can be deselected with »no«. This is useful for certain database servers or test environments, for example. Further details follow in the section»Backup, Recovery«.

Replication

Values: yes, no (standard)
If replication has been selected, the VM and its data are additionally transferred to a 2nd location as a so-called »placeholder«.
This technology makes it possible to get the VM up and running again very quickly, even in the event of a total failure of a site.
Replication is usually only approved for individual systems with priority »red« or »yellow«, as this takes up roughly twice the resources of the actual VM and thus also causes costs that are not small.

Actuality of the deposited data

It is absolutely necessary to keep the data stored in the server database up to date at all times, as this information is the basis for the work of the virtual infrastructure operators. All changes must therefore be registered in advance via the database, even if you make the change yourself (e.g. OS in-place upgrade). Here are the most important points:

  • Information on server operator / deputy,
  • Information on the operating system version,
  • Information on the intended use of the system,
  • IP address / subnet / VLAN information,
  • Classification in test / productive and priority.

If the notification of changes has not been made in advance, it is imperative that a corresponding notification is made immediately after the change has been made.

Management of Virtual Machines

The operating responsibility for a VM begins for the server manager at the level of the operating system upwards. Security settings, updates, configurations, etc. are the responsibility of the server owner(s), as is the management of the applications and services in this VM. This also applies to the up-to-dateness of the pre-installed software, see »Pre-installed Software«.

In principle, virtual servers can be reached via the usual connection paths for remote management. For Windows remote desktop / RDP, for Linux SSH. Extended access is possible via the VMware management interface (vCenter). The login data are:

URL: https://vim-c13.in.uni-weimar.de/ui/
Username: in\<uni-account>    Password: PW of the Uni account

There you have, among other things, the following extended options for managing the VM(s):

  • Direct console access (KVM),
  • Execute Start/Stop/Reset of the VM,
  • Create, manage, delete snapshots,
  • Monitor performance parameters & events,
  • Mount ISO files on a virtual CD/DVD drive

For server administrators who are accessing the vCenter for the first time, it is necessary to have the corresponding firewall rules activated in advance via the user service of the SCC.

Reservation and measures to prevent damage

In exceptional cases, the SCC reserves the right to isolate VMs from the network or switch them off in order to prevent damage to the overall infrastructure of the virtualisation and the Bauhaus-Universität and to be able to maintain the overall operation until clarification. In this case, the person(s) responsible for the server will be contacted immediately to clarify the further procedure. Such cases can be, for example:

  • Outdated operating systems that no longer receive security updates
  • Infection with malware / membership in botnets / sending spam
  • Suspicion of hacking / cyber attacks
  • Unannounced changes in the VM, e.g. wrong IP address set, wrong OS installed
  • Permanently unusual runtime parameters such as 100% CPU load

Usage of Snapshots

With the help of snapshots - snapshots of the VM in its current state - it is possible to freeze a defined state of the VM, e.g. to execute changes, updates, etc. If the attempt fails, the previous state can be quickly restored via the snapshot. If the attempt fails, the previous state can be quickly restored via the snapshot. If the attempt succeeds, the snapshot can be deleted. Snapshots are not a backup, but merely a means of storing an operating state for a short time. It can lead to a number of operational complications if snapshots are too old, too large or too numerous. VMware's recommendations for this are:

https://kb.vmware.com/s/article/1012384?lang=de
https://kb.vmware.com/s/article/1015180?lang=de

In addition to these recommendations, we allow server managers the following maximum values based on many years of operating experience:

  • Max. Number of snapshots per VM: 3 (recommendation: 2-3)
  • Max. Snapshot age: 30 days (recommendation: 1-3)

This means that all snapshots older than 30 days are automatically deleted by the VMware environment. Snapshot chains with more than 3 snapshots are also decimated so that a maximum of the 3 most recent snapshots remain. In total, this means a maximum of 3 snapshots per VM, whereby the oldest is not > 30 days. Every server administrator is also urgently advised to delete snapshots that are no longer needed as quickly as possible; unfortunately, this has been forgotten all too often in the past. A good read on this topic is also the following article:https://cmdb.scc.uni-weimar.de/inventar/download/snapshots_heise_sonderbeilage.pdf

Important notice:
If you use snapshots, the automatic daily backup of the VM cannot be guaranteed. When creating snapshots, please contact the SCC storage group at storage[at]scc.uni-weimar.de immediately.

Pre-installed Software and Settings

Currently, the prefabricated VM templates (base images) are mainly standard minimal installations of the respective operating system with minor adjustments. For example, in Windows the WSUS server for updates is entered and the remote desktop RDP port is set to a different value. There is also the possibility of requesting Microsoft operating systems provided by the ZDM team, these have a whole range of other changes on board; significantly to increase operational security. Efforts are being made to make the VM templates BSI-compliant, but this is not possible at present.

VMware Tools

This software is always pre-installed. The VMware environment communicates and acts with the VMs via the VMware tools. For various things such as starting/stopping the VM and file system operations such as creating snapshots, up-to-date, running VMware tools are essential. The VMware tools for Windows operating systems are provided by VMware, for Linux distributions the open-vmware-tools or open-vm-tools of the respective distribution are to be used. The version must always be kept up to date and an automatic start of the corresponding services must be configured. This software must not be prevented from running.

Malware Protection

Currently, centrally managed malware protection from Sophos is pre-installed on the VMs. For Windows  »Sophos Endpoint Protection« and for Linux »Sophos Protection for Linux«. This software must not be uninstalled or prevented from running. If the VM is operated in a partitioned network, the server administrator must use proxies or similar techniques to create the conditions for regular updates of the signature databases of this software.

Data storage, Partitioning

If possible, the VMs should always be divided into system (operating system, etc.) and data hard disk(s) (databases, web directories, file storage, etc.). VMs with a total of > 200 GB hard disk capacity must be divided in this way. If the capacities of the data exceed 4 TB, a logical volume (LVM) must be created for reasons of performance, backup / recovery and processes for moves / disaster recovery, in which several max. 4 TB virtual hard disks are connected to form a data store. So we do not provide e.g. 10TB large single virtual hard disk. In special use cases it may make sense for the VM to access the underlying storage directly via NFS/RBD/CephFS/SMB/CIFS. This can be implemented in individual cases in consultation with the operators of the virtual infrastructure and the storage group.

Backup, Recovery

All virtual machines with a total disk capacity of up to 1.2 TB are always backed up as an entire VM on a daily basis by default. A different procedure is required for larger VMs. Here, a backup client is installed and configured in the VM so that the large disks are backed up file-based via the client. For VM restore requests, a ticket specifying the host name of the VM and the desired date of the state to be restored must be submitted to the SCC user service. Further information is available under »Data backup«

Important notice:
If you use snapshots, automatic daily backup of the VM cannot be guaranteed. When creating snapshots, please contact the SCC storage group at storage[at]scc.uni-weimar.de.

Monitoring

An integration of systems & servers into the central monitoring (PRTG) is possible on request, an implementation into the database interface of the CMDB is planned. Currently, the person responsible for the server can contact vmware[at]scc.uni-weimar.de by e-mail. The SCC recommends integration for systems with the priority »red« and »yellow«, see paragraph »Priority«.

Special conditions for project servers

The SCC provides such VMs with a maximum duration of one year with the option of extension after prior consultation and approval by the SCC director to support the joint work of research groups as well as the collaboration in projects. An automated workflow for expiration and renewal management is planned, currently the request for renewal is done manually by the SCC. After the first year, costs may be incurred. An overview of our services for such projects can be found here:VMware offerings: Provision of storage capacity or by provisioning virtual servers

Summary of the most important rules

  • New VMs, changes to VMs and deletion of VMs must always be requested via the SCC server database https://cmdb.scc.uni-weimar.de/
  • A responsible person (server operator) and, if possible, a deputy must be named for each VM. This information must always be kept up to date; if necessary, a change request must be made via the SCC server database.
  • Any change within a VM (e.g.: CPU, RAM, disk size, operating system updates) must be requested in advance via the SCC server database, as interventions by the VMware administrator may be necessary.
  • The requested purpose of use must be adhered to; in the event of changes in use, a change request must be submitted in advance via the SCC server database (form field Server function)
  • Management interface of the virtual machines accessible via a browser / usable at https://vim-c13.in.uni-weimar.de or https://vim-s6a.un-weimar.de
  • The following pre-installed software must not be removed or prevented from functioning:
    • VMware Tools (resp. open-vm-tools)
    • Malware protection
  • The malware protection must always be up-to-date (automatic update) and activated
  • Except for Windows, the open-vm-tools of the distribution provider are usually to be used, not the VMware tools from VMware
  • Rules for securing operations via snapshots:
    • max. 3 snapshots per VM are allowed
    • max. age of a snapshot 30 days
  • Furthermore, the rules of the SCC backup/restore policy under »Data backup« apply
  • If publicly visible web content is provided via the VM, the guidelines for websites / microsites must be observed
  • The SCC reserves the right to switch VMs offline in exceptional cases in order to avert possible dangers and to be able to maintain the overall operation until clarification.

Support and contact

If you have any questions about the virtual infrastructure at SCC, please contact the Infrastructure Operators team at vmware[at]scc.uni-weimar.de.

Supplementary documents and websites