2
0
mirror of https://github.com/xcat2/confluent.git synced 2026-09-29 08:41:00 +00:00

Compare commits

...

326 Commits

Author SHA1 Message Date
Jarrod Johnson df7cba00fd Amend the message on collective failure 2018-08-17 16:45:45 -04:00
Jarrod Johnson dfb720d0ee Have collective command warn if the libssl library is not viable
Main example is RedHat providing pyOpenSSL of relatively ancient
vintage.
2018-08-17 13:57:13 -04:00
Jarrod Johnson 9b48110155 Do not proceed a logged, but broken session
It shouldn't be possible for this to be the case, but out of an
abundance of caution, check for this.  So far only produced this by
forcing broken = True in a debug session.  Intended to catch an alleged
scenario where console was managing to use a broken session (fixed in
pyghmi) and have confluent also recognize the situation for non-console
usage).
2018-08-16 14:43:16 -04:00
Jarrod Johnson 3064e7bef6 Ensure path is made prior to creating transactioncount
Fresh install will be missing /etc/confluent/cfg.  Advance the
_mkpath call to fix this problem.
2018-08-08 18:05:45 -04:00
Jarrod Johnson 1d4df8af3a Fix extraneous error in log on connectivity loss 2018-08-07 15:43:53 -04:00
Jarrod Johnson 2aba6e469c Correct variable name in the 'connected' fix 2018-08-07 15:31:41 -04:00
Jarrod Johnson de58593f14 Fix inability to notice underlying broken layers of the SOL
Through an unknown set of circumstances, an solconnection could be
stuck 'connecting'.  In every case analyzed, the ipmi_session was
broken.  Use that to detect a class of failure and react appropriately.
2018-08-07 15:12:53 -04:00
Jarrod Johnson ecbe1a86b1 Revert "Have nodeconsole restore term on exit"
This reverts commit 2972374da8.
2018-08-02 10:27:37 -04:00
Jarrod Johnson 2972374da8 Have nodeconsole restore term on exit 2018-08-02 10:07:41 -04:00
Jarrod Johnson 81dd6202d3 Fix when rpc has no 'exc' but has 'xid' 2018-07-30 11:26:09 -04:00
Jarrod Johnson 36a202842a Fix collective on rpc exception
Exceptions on collective calls were not correctly handled, fix
the handling so that collective continues and also the calling function
is correctly given the exception.
2018-07-30 09:33:24 -04:00
Jarrod Johnson 6a8e24dd0e Prioritize interactive feedback part of console handling. 2018-07-26 08:55:25 -04:00
Jarrod Johnson d3afeb3414 Fix web shell if user hits enter too fast 2018-07-24 17:20:22 -04:00
Jarrod Johnson 1bf4c0ac0a Have collective coalesce watched updates
Particularly chatty output can make collate be unreasonable in
low quality terminals and links.  Throttle to about 4 times a second.
2018-07-24 16:50:46 -04:00
Jarrod Johnson 8e422ef822 Fix ssh access
Fixed handler (e.g. ssh) did not return console consistent with
the plugin defined handlers.
2018-07-24 16:48:46 -04:00
Jarrod Johnson f0edbbad39 Have collective show present some info when not in quorum 2018-07-20 14:11:38 -04:00
Jarrod Johnson 5cf1671350 Make the takeover process more deterministic
Try to avoid submitting to be a follower while we are currently
becoming a leader
2018-07-20 13:50:42 -04:00
Jarrod Johnson e5c4219ee9 Reorder certificate check
First order of business is to verify certificate before even thinking
about if the request is possible
2018-07-20 13:34:14 -04:00
Jarrod Johnson 3ff7e42074 Change behavior for fallback handling
Fallback would do nothing to fix a persistent problem with an IPMI
session.  For lack of knowing how to avoid the situation, at least
make changes so it won't go wrong in the future.
2018-07-20 13:20:50 -04:00
Jarrod Johnson fab177e077 Fix node[group][attrib|define] handling of =
Attributes with = in the value were not handled correctly,
fix by only doing one split.
2018-07-20 09:54:17 -04:00
Jarrod Johnson a1ba5f59a8 Fix collective show on non-collective 2018-07-19 17:21:01 -04:00
Jarrod Johnson 9bcca6bfad Provide collective show on all members 2018-07-19 17:08:20 -04:00
Jarrod Johnson 96671ace4e Correct collective show behavior 2018-07-19 16:48:30 -04:00
Jarrod Johnson bcff3fc962 Improve collective show readability 2018-07-19 16:39:13 -04:00
Jarrod Johnson 54d93571d1 Have leader provide more data in collective show 2018-07-19 16:26:05 -04:00
Jarrod Johnson f2f902de7b Have collective show report when collective inactive
Collective show was misleading if not in a collective.
2018-07-19 15:59:15 -04:00
Jarrod Johnson a09792f969 Schedule periodic attempts to restart collective
If collective is lost due to connectivity, this will cause
occasional attempts to bring it back.
2018-07-19 15:49:05 -04:00
Jarrod Johnson 7d16c943a8 Handle updating address of collective member on connect
If a collective member changes its IP address, update at the next
possible opportunity.
2018-07-19 15:24:08 -04:00
Jarrod Johnson b053d41cd8 Error on loss of manager in flight 2018-07-19 14:36:23 -04:00
Jarrod Johnson 200569e7af Merge branch 'master' into clustertime 2018-07-19 13:32:00 -04:00
Jarrod Johnson c3c0e1570a Push quorum state to followers
The followers need to know quorum state.
2018-07-19 13:27:21 -04:00
Jarrod Johnson 10c82a72b5 Restore message on unreachable collective member
The parallel execution had broken how that message transmits.

Bonus, make it a per node error.
2018-07-18 16:49:54 -04:00
Jarrod Johnson 79cdf65a72 Fix SLES sockapi
Previous fix was applied to the incorrect section of code
2018-07-18 15:07:22 -04:00
Jarrod Johnson 497ca40492 Do not abort connecting process on bad cert
The target may be non-viable, but don't let that ruin the party
for everyone.  Let it keep going as if the system were down.
2018-07-18 14:58:16 -04:00
Jarrod Johnson fd33e6ae01 Fix non-collective confluent mode
list_collective returns an iterator, which will be True...
2018-07-18 14:53:23 -04:00
Jarrod Johnson 32f944e67c Handle unclean loss of current proxy host
If transition is less than gentle, provide a path to restore automatic
if it gets moved.
2018-07-18 14:32:39 -04:00
Jarrod Johnson dcad9f5a75 Add keepalive and acks to collective
Detect unplugged condition (eventually).
2018-07-18 13:45:03 -04:00
Jarrod Johnson 2a34388d09 Add -p to man page for nodepower 2018-07-18 11:02:12 -04:00
Jarrod Johnson 6993e0b496 Fix nodepower argument parsing
nodepower was assuming that the second parameter was always the
state regardless of option parsing.  Use args instead to fix.
2018-07-18 11:00:01 -04:00
Jarrod Johnson b7fe72673d Add clear node/group attributes to collective
collective was not syncing clear directives.
2018-07-17 15:57:48 -04:00
Jarrod Johnson 0159bf1b1d Fix typo in error message 2018-07-17 15:39:08 -04:00
Jarrod Johnson cf9ad11290 Short out operations if in collective mode but no collective.manager 2018-07-17 15:25:12 -04:00
Jarrod Johnson ddd7ef5eba Fix proxyconsole break and reopen 2018-07-17 15:05:09 -04:00
Jarrod Johnson 73da8ec8b5 Fix ProxyConsole if self.remote is not yet set 2018-07-17 14:44:59 -04:00
Jarrod Johnson eac4d97732 Disengage remote console on manager change
This results in a more direct treatment of manager change.
2018-07-17 13:10:01 -04:00
Jarrod Johnson fa9ecfbb94 Merge branch 'clustertime' of github.com:jjohnson42/confluent into clustertime 2018-07-17 11:46:53 -04:00
Jarrod Johnson fc5472065a Catch missing '@' in token as invalid token 2018-07-17 11:46:40 -04:00
Jarrod Johnson cb0845596e Provide explanation about nodemedia list and no media. 2018-07-17 11:20:27 -04:00
Jarrod Johnson 0d936e0059 Ensure no more than one in-flight slave connection from a given follower
This will prevent a connection from deregistering itself after the
replacement registers itself.
2018-07-17 10:36:31 -04:00
Jarrod Johnson a7b8f0ab0c Parallelize cross-manager requests
Rather than doing it at one at a time, parallelize the requests
for improved performance.
2018-07-17 10:07:32 -04:00
Jarrod Johnson 3ab4203104 Explicitly set ECDHE curve
Some vintages of the SSL stack require we explicitly request a curve,
so here it is.
2018-07-16 16:23:33 -04:00
Jarrod Johnson 13aa2e9aae Catch more broad errors
Operating on a closed socket is not a socket.error
2018-07-16 11:58:18 -04:00
Jarrod Johnson 7462bc28e8 Use the eventlet socket in configmanager 2018-07-16 10:06:53 -04:00
Jarrod Johnson 18f1c07d65 Change to setting an errstr rather than exception
If nodefirmware update has an issue, provide error message instead.
2018-07-16 09:03:02 -04:00
Jarrod Johnson 0016077bee Ensure that wait_for_sync always does a new sync
If a sync is in progress, wait for that to complete.

Then issue the requested *new* sync.

Probably only needed if fullsync, as the one in progress may be a
'dirty' only sync and fullsync would be satisfied by the partial sync
without it, which is bad.
2018-07-13 22:15:38 -04:00
Jarrod Johnson 1dad69097b Be consistent with sync during load of leader cfg
Pass through sync as appropriate.

Also changes meant for previous commit
2018-07-13 21:52:17 -04:00
Jarrod Johnson fd7c428d1f Cleanup leftover sockets and more reliably be following or leading
Before there was a chance to be in a half state, leading to an inability
to reach consensus on leader.
2018-07-13 21:20:42 -04:00
Jarrod Johnson 80a1bd72e7 Correct arguments for Thread constructor 2018-07-13 15:43:09 -04:00
Jarrod Johnson 042d7ab5cf Modify clear_commit to use the same thread
Additionally, wrap a lock around the dbm operations, in case something
in the future makes a mistake.
2018-07-13 15:27:16 -04:00
Jarrod Johnson c74fdf5924 More collective join errors 2018-07-13 11:07:39 -04:00
Jarrod Johnson 58bf226d23 Relay error from server about token issue 2018-07-13 10:50:17 -04:00
Jarrod Johnson 6f012b69a1 Provide cleaner message for collective manager being unreachable 2018-07-13 10:43:20 -04:00
Jarrod Johnson 7f1e5d2302 Add explanation of 'all' in nodeattrib man page 2018-07-13 09:57:23 -04:00
Jarrod Johnson 3e2a827ff9 Correct typo in nodeattrib man page 2018-07-13 09:50:08 -04:00
Jarrod Johnson 1d16534c16 If replacing a follower stream, ensure the old one closes 2018-07-13 09:37:00 -04:00
Jarrod Johnson c80ebb0e8d Explicitly close connection before replacement
If an existing follower is stalled out, close the socket explicitly
to avoid leaving it open in lsof.
2018-07-13 09:14:36 -04:00
Jarrod Johnson efaf1dae70 Make cfgleader modifications more robust
If cfgleader is about to forget a socket, explicitly try to close
it first.
2018-07-13 09:05:28 -04:00
Jarrod Johnson 1de82936ed Add full sync mode
For implementing clear config, all data must be presumed dirty.
2018-07-12 17:06:37 -04:00
Jarrod Johnson b0c384c9ca Check quorum on attribute read
It's too bizarre for attribute read from api to work
without quorum, could be misleading.
2018-07-12 16:05:04 -04:00
Jarrod Johnson 61dd71778f Never generate new key on crypt read
An autogenerated key on read can never be useful.  Instead, let it fail
and assume a repair action is coming.
2018-07-12 15:55:05 -04:00
Jarrod Johnson 0f3014957b Fix non-ascii unicode handling of consoles 2018-07-12 14:16:44 -04:00
Jarrod Johnson c925353f02 Fix rollback
First the data was not actually being staged in the rollback area.
Secondly, there wasn't an assurance that the disk wouldn't have rollback
committed...
2018-07-12 10:01:32 -04:00
Jarrod Johnson 9edb225bd3 Fix the reference to sync to file 2018-07-12 09:20:41 -04:00
Jarrod Johnson 7cdc3c1400 Implement clear config rollback
Should something go awry during config
load, rollback the clear and load.
2018-07-12 08:48:21 -04:00
Jarrod Johnson bd2a3f14e6 Defer txcount increment until after potential failure
This avoids the txcount incrementing when no transaction actually
occurs.
2018-07-12 08:41:07 -04:00
Jarrod Johnson 87d00b7447 Fix typo in manpage for nodeattrib 2018-07-11 16:48:41 -04:00
Jarrod Johnson beedfb0600 If a drone doesn't exist, treat it as if it's an invalid certificate 2018-07-11 16:29:45 -04:00
Jarrod Johnson ce59a36351 Avoid excessive syncs on connect
This removes some redundancy and avoids writing and loading to disk
during the initialization process.
2018-07-11 16:07:56 -04:00
Jarrod Johnson a1e612968e Merge branch 'clustertime' of github.com:jjohnson42/confluent into clustertime 2018-07-11 13:44:30 -04:00
Jarrod Johnson ce1a58bf58 Provide a more specific error when file doesn't exist
nodemedia does check locally, but server should also check it, in
case of remote *and* service nodes.
2018-07-11 13:17:37 -04:00
Jarrod Johnson b386084f2d Note future enhancement for flattening on collective manager change... 2018-07-11 11:07:59 -04:00
Jarrod Johnson 8e9bcbb44f Clear txcount on enroll
The transaction count on 'join' was being honored as high, when
it never should be.
2018-07-11 09:40:22 -04:00
Jarrod Johnson 704aaeecf9 Tolerate newline in myname
vim is quite insistent on adding a newline, tolerate that.
2018-07-11 09:36:51 -04:00
Jarrod Johnson 11968faffc Numerous fixes to collective
If client has higher transaction count, do not close the connection
before extracting peer address.

If our connect session is rudely terminated, abort rather than trying
to continue.

On assimilate failure, ignore a failed assimilate with no data.

Fix problem where a follower getting double deleted was causing an error.
2018-07-10 14:55:57 -04:00
Jarrod Johnson 8769c438c0 Relax cryptodomex requirement
We don't *need* the higher version, it just accelerates password auth
if possible.  Fix conflict with RH provided package.
2018-07-10 10:10:04 -04:00
Jarrod Johnson f1489bf527 Add all to SYNOPSIS of nodeattrib 2018-07-09 16:52:07 -04:00
Jarrod Johnson c03781c022 Add 'all' to usage message of nodeattrib 2018-07-09 16:49:45 -04:00
Jarrod Johnson 298e11f60f Allow invite from non-leader role
A non-leader transaction is modified such that the enroll node
can be connected to the leader and have validation.
2018-07-09 16:40:43 -04:00
Jarrod Johnson 67d6e9a6c7 Add collective show
Provide a harmless way to look at collective state
2018-07-09 15:07:24 -04:00
Jarrod Johnson a4edf9afb8 Rename confluentutil to collective
Also adjust output to be a bit more automation friendly.
2018-07-09 13:33:56 -04:00
Jarrod Johnson 2342fe717e Remove superfluous call to sync to file
load_from_json already makes the call, remove the extra call that is
redundant.
2018-07-09 12:59:37 -04:00
Jarrod Johnson 08cf698609 Only conditionally require ffi
Only collective mode requires ffi, do not incur requirement for
non-collective mode.
2018-07-09 09:41:10 -04:00
Jarrod Johnson a905ae4865 Only zap cfgleader if not set
stale relay_slaved_requests was trouncing following status, correct
by avoiding zapping that state
2018-07-03 14:43:23 -04:00
Jarrod Johnson 1eaf5357ca Resolve race conditions on simultaneous collective outage
Implement random backoff strategy for serializing connect out and
connect in.
2018-07-03 14:09:09 -04:00
Jarrod Johnson 39378170b1 Add an error if cfgstreams loses a follower 2018-07-03 11:13:23 -04:00
Jarrod Johnson 2156e99ae4 Use alternative name for pycryptodomex 2018-07-03 10:43:05 -04:00
Jarrod Johnson bba12ed9e7 Update requirements to pull in cryptodomex 2018-07-03 10:32:33 -04:00
Jarrod Johnson 282043ed97 Switch to cryptodome
Cryptodome is a modern, but compatible replacement for pycrypto.

We may move to cryptography eventually, but start with this for now
for some nice speedups in some cases.
2018-07-03 10:31:13 -04:00
Jarrod Johnson 09c239b294 Fix memory leaks
For one, configmanager was left with stale callback references, clean
those up.

For another, the callback pattern was creating a circular reference that
python memory management couldn't overcome.  Break the reference
explicity when an item is disposed of.
2018-07-03 08:58:32 -04:00
Jarrod Johnson 0b26d12837 Fix memory leaks
For one, configmanager was left with stale callback references, clean
those up.

For another, the callback pattern was creating a circular reference that
python memory management couldn't overcome.  Break the reference
explicity when an item is disposed of.
2018-07-03 08:58:11 -04:00
Jarrod Johnson 956faee052 Correct typo in variable name 2018-06-28 14:12:56 -04:00
Jarrod Johnson 8e77466875 Carry txcount through connect replica
The txcount was not up to date when offline updates occurred.
2018-06-28 13:54:31 -04:00
Jarrod Johnson f01c7e19f7 Fix HA healing problems
If we are superior to leader, become leader
When abandoned by a leader, forget the leader.
2018-06-28 13:51:52 -04:00
Jarrod Johnson 5d894912ac Clean up more potential for CLOSE_WAIT
Explicitly close sockets in a few more places.
2018-06-28 10:50:03 -04:00
Jarrod Johnson 93e78d1f03 Fix try_assimilate
Also get rid of CLOSE_WAIT in some situations.
2018-06-28 09:53:34 -04:00
Jarrod Johnson 97f932b946 Fix the attempt to assimilate
In the wake of leader loss, the assimilate attempt was causing one
of a couple of bad behaviors, a trace if broken and if it should
work, the format of the data payload was incorrect.
2018-06-26 15:29:51 -04:00
Jarrod Johnson 61f24f4d8a Fix mismatched braces in previous commit 2018-06-26 15:09:17 -04:00
Jarrod Johnson 68a79695fc Add check for quorum to invite generation process
If a would-be collective member has no quorum, provide a clear
message indicating this as an issue.
2018-06-26 14:55:42 -04:00
Jarrod Johnson 401352998c Correctly show the error on non-leader
When non-leader tries to invite, print the error rather than unhelpful
exception with no helpful data.
2018-06-26 14:35:23 -04:00
Jarrod Johnson 9eeede651e Error when inviting from non leader 2018-06-26 14:24:52 -04:00
Jarrod Johnson fb98dd5636 Fix packaging issues 2018-06-26 14:12:53 -04:00
Jarrod Johnson 0843b991ea Isolate following to dedicated greenthread and put fallback to assert leadership 2018-06-26 14:06:01 -04:00
Jarrod Johnson 11e6145a46 Add quorum check to non-config requests 2018-06-26 13:52:20 -04:00
Jarrod Johnson 7433dd3e38 Wrap cfg init on follow in lock
Use a lock to provide more atomic behavior for connecting should
something go wrong in calling connect_to_leader incorrectly.
2018-06-26 10:13:27 -04:00
Jarrod Johnson 4de1fea7aa Implement attempt to assimilate 2018-06-25 17:30:29 -04:00
Jarrod Johnson f6342dd31f Allow connect_to_leader to cycle in the parent for loop
Startup was foiled when one entry was bad.  Also add comments
on invite/join needing to be done from current leader, and the need
for better error messages.
2018-06-25 14:51:39 -04:00
Jarrod Johnson 0499a14fe8 Try to reestablish collective on leader loss 2018-06-25 10:40:29 -04:00
Jarrod Johnson f301cd272f Have follower notice lost leader
This will mitigate period during which requests
hang instead of error.
2018-06-24 11:08:15 -04:00
Jarrod Johnson bf4f5ad5ae Recognize loss of follower as step toward loss of quorum
Properly reap the loss of a follower.
2018-06-22 16:17:48 -04:00
Jarrod Johnson 5a5f0169a7 Fix logic on null messages 2018-06-22 15:46:41 -04:00
Jarrod Johnson 03adec089d Also clear cfgleader when actually becoming leader 2018-06-22 15:26:08 -04:00
Jarrod Johnson f065fafd61 Clear the quorum block on initial xfer
The initial transfer needs to have the configmanager working locally.
2018-06-22 15:23:20 -04:00
Jarrod Johnson b315042850 Error if starting without quorum
If collective mode is present, but no candidates worked, still error
out.
2018-06-22 15:12:42 -04:00
Jarrod Johnson 5588712320 Clear transaction count on config clear
A clear configuration should have 0 transactions.
2018-06-22 14:46:52 -04:00
Jarrod Johnson c6a0aeca3b Fix dispatch of commands with InputData
Inputdata needed to be serialized for the network.  Further, had
to have a JSON-safe payload for indicating name for certificate look
up, to avoid doing pickle load on client input prior to client
validation.
2018-06-22 14:41:41 -04:00
Jarrod Johnson f45228c067 Auto-migrate open console session on loss of management
This will keep consoles connected across a manager change.  Note
that as it stands, a console session enduring multiple changes will
become increasingly convoluted as the redirect chain will increase
every time, rather than shorting back to the entry point.
2018-06-21 15:14:56 -04:00
Jarrod Johnson d34e65f9b7 Fix tlvdata handling of unicode input
Unicode input is normalized to bytes.  Also have to handle python3
not having 'unicode', do a quick change to support that in both.
2018-06-21 14:29:54 -04:00
Jarrod Johnson baef8c2c42 Connect and fix the proxyconsole
This works full duplex for sockapi and read for web consoles.
2018-06-21 13:46:51 -04:00
Jarrod Johnson 1543e145b7 Draft for proxyconsole object for remote use of consoles
This would be the stub stand in for the console object to
connect to remote console object rather than local.
2018-06-20 16:36:59 -04:00
Jarrod Johnson 2d403f7d68 Defer import of confluent log
confluent.util pulls in log more cleanly, for now rearrange for
easier time running slp test mode.
2018-06-20 14:56:07 -04:00
Jarrod Johnson 78117a1b1a Draft proxyconsole support for sockapi
Foundation for consoleserver to be able to do backend to backend
connections
2018-06-19 16:50:49 -04:00
Jarrod Johnson 624012774c Discontinue processing on cert mismatch
If cert mismatched, the client was shut out, but processing continued,
fix this mistake.
2018-06-19 16:45:06 -04:00
Jarrod Johnson 003358da5f Discontinue ipmi session on collective manager change 2018-06-19 15:10:59 -04:00
Jarrod Johnson 6b24b9691b Merge branch 'master' into clustertime 2018-06-19 11:07:53 -04:00
Jarrod Johnson 6ba8ca2fa2 Remove accidental change
keepalive was disabled, which negatively impacted
web ui performance.  Re-enable.
2018-06-19 11:07:23 -04:00
Jarrod Johnson 38898ca921 Auto-make certificate if missing
Automatically fix a missing certificate if this is the case.
2018-06-19 11:05:38 -04:00
Jarrod Johnson 810be71720 Initial support for non-console dispatch
For non-exceptional cases, it is now functional.
2018-06-15 15:54:26 -04:00
Jarrod Johnson b877a95645 Include absent devices in the json of nodeinventory 2018-06-15 11:03:03 -04:00
Jarrod Johnson efd5732682 Amend json output
Have the nodeinventory json output in a bit more directly useful format,
rather than regarding the API structured JSON...
2018-06-15 11:02:57 -04:00
Jarrod Johnson 4906d6e9c4 Add --json to nodeinventory
Have nodeinventory have an option to output in json.
2018-06-15 11:02:51 -04:00
Jarrod Johnson f2500d9d27 Add general confluentutil command
This provides util commands to manage certificates and collective
membership.
2018-06-13 16:23:49 -04:00
Jarrod Johnson 0507e89da8 Add ability to skip key backup and interactive password
Backups should carefully protect keys.json, but that's only feasible
interactively.  However keys don't change, so have a way to combine
protected keys.json with password with relatively safe non-interactive
incremental backups.
2018-06-13 16:22:40 -04:00
Jarrod Johnson 8cea5a4fed Fix redacted db dump
The db dump function did not trigger the init() function.
Rearrange the collective section to occur after a section that does
ensure the configmanager is initted.
2018-06-13 09:15:27 -04:00
Jarrod Johnson 8515d43dad A shell script to illustrate generating ECDSA key
For now put down the openssl commands required to get the key and
certificate available.
2018-06-12 16:57:36 -04:00
Jarrod Johnson 5c12dc2cba Do not require exactly TLSv1.0
This was breaking TLSv1.2.
2018-06-08 10:15:38 -04:00
Jarrod Johnson a08eace00c Fix missing import of traceback 2018-06-05 13:08:05 -04:00
Jarrod Johnson cdb20c0302 Fix missing import of traceback 2018-06-05 13:05:42 -04:00
Jarrod Johnson 3eed46ec08 Disconnect on loss of ownership 2018-06-02 12:36:20 -04:00
Jarrod Johnson daef9fa60b Fix confusing nodeconfig error handling
Properly react to error conditions
2018-06-01 16:48:34 -04:00
Jarrod Johnson a7a4ede580 Fix confusing nodeconfig error handling
Properly react to error conditions
2018-06-01 16:48:19 -04:00
Jarrod Johnson a560dc1974 Add timeout on httpapi socket
Clients that fail to send any data, or keep a persistent socket
open without using it are killed off.
2018-06-01 16:26:21 -04:00
Jarrod Johnson e8cea66a85 Add timeout on httpapi socket
Clients that fail to send any data, or keep a persistent socket
open without using it are killed off.
2018-06-01 16:25:38 -04:00
Jarrod Johnson 8fc29d1b46 Correctly allow manual discovery through ambiguous situation
Manual discovery may 'catch' some incidental data from auto-discovery.
Since the operation is manual, trust the user rather than assume the
user is confused.
2018-06-01 16:09:32 -04:00
Jarrod Johnson 6e0186947f Correctly allow manual discovery through ambiguous situation
Manual discovery may 'catch' some incidental data from auto-discovery.
Since the operation is manual, trust the user rather than assume the
user is confused.
2018-06-01 16:09:13 -04:00
Jarrod Johnson a45a0e31f1 Merge branch 'master' of github.com:jjohnson42/confluent 2018-06-01 15:54:31 -04:00
Jarrod Johnson 8ea532053e Add guardrail around slp snoop
If a program error should befall our poor slp service, log the issue
and carry on.
2018-06-01 15:54:27 -04:00
Jarrod Johnson 366c955956 Add guardrail around slp snoop
If a program error should befall our poor slp service, log the issue
and carry on.
2018-06-01 15:53:52 -04:00
Jarrod Johnson d968fb6b04 Add TODO note on how to do the collective console sessions
We will hove consoleserver abstract this away from sockapi/httpapi
concerns.
2018-05-31 16:22:11 -04:00
Jarrod Johnson db1d5d6dff Prepare for request dispatch
Sort and exclude nodes that have collective.manager, prepare to relay
for dispatch.
2018-05-31 16:18:58 -04:00
Jarrod Johnson 867dd0dda7 Handle collective.manager being set back to the node 2018-05-31 16:16:03 -04:00
Jarrod Johnson 22b63df036 First pass at suppressing foreign consoles
If console session is another node, prevent it from working locally.
Focus first on making sure the wrong nodes don't work before routing.
2018-05-31 15:53:09 -04:00
Jarrod Johnson 54f419d9b4 Add 'collective.manager'
For requests that must be routed to a definitive owner (e.g. IPMI),
have an attribute to track the currently responsible server for
singleton operations.
2018-05-31 15:47:19 -04:00
Jarrod Johnson 5d6e241287 Fix deadlock on collective create user
There was a call to rpc-out from within a 'true' function, which is bad,
fix this by only calling the 'true' function from 'true' function.
2018-05-30 16:21:40 -04:00
Jarrod Johnson 417bc5acda Fix incorrect function call for user creation
In collective mode, the incorrect rpc call was made
2018-05-30 15:30:27 -04:00
Jarrod Johnson 57ffe166d3 Merge branch 'master' into clustertime 2018-05-25 10:24:34 -04:00
Jarrod Johnson 5718c60a51 Merge branch 'clustertime' of github.com:jjohnson42/confluent into clustertime 2018-05-25 10:23:25 -04:00
Jarrod Johnson 31effcc025 Fix mistake in variable name in nodeconfig 2018-05-25 10:22:35 -04:00
Jarrod Johnson cefca49128 Fix mistake in variable name in nodeconfig 2018-05-25 10:21:34 -04:00
Jarrod Johnson 41a5eaa464 Merge branch 'clustertime' of github.com:jjohnson42/confluent into clustertime 2018-05-23 13:02:12 -04:00
Jarrod Johnson 46c4065b81 Correct another typo 2018-05-23 13:02:03 -04:00
Jarrod Johnson 57e323786e Fix syntax error in the recent code 2018-05-22 11:00:30 -04:00
Jarrod Johnson 3b2a18a650 Merge branch 'master' into clustertime 2018-05-22 10:34:31 -04:00
Jarrod Johnson 3ace7747ab Fix typo 2018-05-22 10:11:37 -04:00
Jarrod Johnson 8ede0fd8ef Document {{}} escape on noderun and nodeshell
Documentation did not explain that
2018-05-22 09:59:30 -04:00
Jarrod Johnson be3ecf60a5 Fix bad error message on {} in nodeshell/noderun
{} used in awk is likely, give proper error message.
2018-05-22 09:56:53 -04:00
Jarrod Johnson 8807fcfd22 Fix missing portname in lldp data
Root cause was pysnmp returning extraneous leftover data causing
calling code to overrite good data.
2018-05-22 09:37:01 -04:00
Jarrod Johnson 33fe0a3db4 Fix wrong port name for G8332
Was using the incorrect half of the return, which broke on G8332.
2018-05-22 09:36:55 -04:00
Jarrod Johnson ca7711b373 Fix missing portname in lldp data
Root cause was pysnmp returning extraneous leftover data causing
calling code to overrite good data.
2018-05-22 09:36:09 -04:00
Jarrod Johnson 8b37199654 Fix wrong port name for G8332
Was using the incorrect half of the return, which broke on G8332.
2018-05-22 09:34:02 -04:00
Jarrod Johnson ff2bc89fae Merge branch 'master' into clustertime 2018-05-21 15:53:48 -04:00
Jarrod Johnson 5fe2d2a31c Fix unprintable characters in some chassisid
Some switches send raw octets back, some printable.  Try to normalize
when unprintable chassis id are detected.  This is not 100%, if the hex
would be all between 20 and 80 throughout the string, then this will
fail to do the right thing.

Hopefully, the amount of times when lldp partners disagree on how to
implement LLDP-MIB will be limited.  Currently it is known than Lenovo
and Juniper switches disagree, and both of those have what would
be unprintable values in the mfg portion of the chassis id.
2018-05-21 15:53:42 -04:00
Jarrod Johnson a4fed0601c Fix unprintable characters in some chassisid
Some switches send raw octets back, some printable.  Try to normalize
when unprintable chassis id are detected.  This is not 100%, if the hex
would be all between 20 and 80 throughout the string, then this will
fail to do the right thing.

Hopefully, the amount of times when lldp partners disagree on how to
implement LLDP-MIB will be limited.  Currently it is known than Lenovo
and Juniper switches disagree, and both of those have what would
be unprintable values in the mfg portion of the chassis id.
2018-05-21 15:53:12 -04:00
Jarrod Johnson 41298a8e01 Extend collective data functions to more functions
Add to users and groups.  Refactor reusable code.
Code that remains still looks awfully repetitive though...
2018-05-21 15:46:51 -04:00
Jarrod Johnson caa4000b7e Merge branch 'master' into clustertime 2018-05-21 11:49:19 -04:00
Jarrod Johnson fde2c7a8e0 Fix the encuuid reference
encuuid is a list, not the value, so get the first value
rather than try to concatenate the string.
2018-05-18 11:49:06 -04:00
Jarrod Johnson fbbb5d048f Fix the encuuid reference
encuuid is a list, not the value, so get the first value
rather than try to concatenate the string.
2018-05-18 11:47:34 -04:00
Jarrod Johnson 32d60145f7 Fix typo in discovery core 2018-05-18 10:20:57 -04:00
Jarrod Johnson 1db781852c Fix typo in discovery core 2018-05-18 10:20:31 -04:00
Jarrod Johnson 0dbf82b0f1 Clean up errors on bad ipv4 addresses
confluent errors are better now
2018-05-17 16:24:31 -04:00
Jarrod Johnson 675dc966c7 Clean up errors on bad ipv4 addresses
confluent errors are better now
2018-05-17 16:24:06 -04:00
Jarrod Johnson a9485706d1 Update warning to be commented out, just in case.. 2018-05-17 15:40:59 -04:00
Jarrod Johnson db1ae03415 Sample script for mac to ipv6 translation
Useful for some generic applications where nodediscover
does not have full support, but must be used with care
as it doesn't guarantee the mac address is what we expect
it to be.
2018-05-17 15:40:49 -04:00
Jarrod Johnson 9826235d4d Update warning to be commented out, just in case.. 2018-05-17 15:40:20 -04:00
Jarrod Johnson 232140899e Sample script for mac to ipv6 translation
Useful for some generic applications where nodediscover
does not have full support, but must be used with care
as it doesn't guarantee the mac address is what we expect
it to be.
2018-05-17 15:35:52 -04:00
Jarrod Johnson 5dddae0ebf Cleaner handling of invalid names in restore attempt
Detect problems ahead af time and more cleanly print a message.
2018-05-17 14:40:40 -04:00
Jarrod Johnson 39e9bf0be5 Cleaner handling of invalid names in restore attempt
Detect problems ahead af time and more cleanly print a message.
2018-05-17 14:40:19 -04:00
Jarrod Johnson d6b7c536d5 Fix discovery of old SMM firmware
Older SMM firmware will not have neighbor data, ignore and move on
in such a case.
2018-05-17 14:21:24 -04:00
Jarrod Johnson f21db46cdd Fix discovery of old SMM firmware
Older SMM firmware will not have neighbor data, ignore and move on
in such a case.
2018-05-17 14:20:59 -04:00
Jarrod Johnson 2d1ba7cc9b Merge branch 'master' into clustertime 2018-05-17 13:13:46 -04:00
Jarrod Johnson 22049002bb Fix exitcode references before use 2018-05-17 11:11:11 -04:00
Jarrod Johnson 727aa9a56c Merge branch 'clustertime' of github.com:jjohnson42/confluent into clustertime 2018-05-16 11:28:08 -04:00
Jarrod Johnson dcb1c2b32b Fix load of txcount
Mistake caused txcount not to restore from disk.
2018-05-16 11:27:46 -04:00
Jarrod Johnson d705d6320a Start setting the stage for leader change on restart
Have connect() have a way to recover if leader is dead.

Also these will be involved in configmanager detected loss of leader
2018-05-16 11:27:46 -04:00
Jarrod Johnson 52e2038fdf Fix transaction count in collective
Slave members were not persisting the value to disk
2018-05-16 11:27:46 -04:00
Jarrod Johnson 9d58a2d382 Correct scope of currentleader 2018-05-16 11:27:46 -04:00
Jarrod Johnson a2187087f7 Fix not having currentleader set
A slave node now recognizes itself as such.
2018-05-16 11:27:46 -04:00
Jarrod Johnson c5b5178f39 Block some early startup problems in collective 2018-05-16 11:27:46 -04:00
Jarrod Johnson ff026ee034 Include absent devices in the json of nodeinventory 2018-05-16 11:27:46 -04:00
Jarrod Johnson 1cc659a3b0 Amend json output
Have the nodeinventory json output in a bit more directly useful format,
rather than regarding the API structured JSON...
2018-05-16 11:27:46 -04:00
Jarrod Johnson 8bc8faf0bc Add --json to nodeinventory
Have nodeinventory have an option to output in json.
2018-05-16 11:27:46 -04:00
Jarrod Johnson 1c2c9931a8 Persist the transactioncount
Needed for eventually ascertaining the viability in selecting leader.
2018-05-16 11:27:46 -04:00
Jarrod Johnson 778a153170 Correct spelling error 2018-05-16 11:27:46 -04:00
Jarrod Johnson 34c510e30a Try to persist name as myname
hostname may not agree with the name chosen by user, in such a case
persist the name and use that, falling back to gethostname()
as needed.
2018-05-16 11:27:46 -04:00
Jarrod Johnson d4babbffa4 Check and try to start collective on startup
Not yet good enough for a leader to rejoin, but enough for a follower
to rejoin automatically.
2018-05-16 11:27:46 -04:00
Jarrod Johnson 1c930eba9d Have attrib set wait on all collective members
This will mean that it is reliable that a nodeattrib ; <command>
in delegation scenarios is guaranteed to execute in order.
2018-05-16 11:27:46 -04:00
Jarrod Johnson c4b564123f Allow slave collective drones to set
It works (once), but needs xid fix.
2018-05-16 11:27:46 -04:00
Jarrod Johnson 6d728df4dc Fix reuse of channel for receiving changes 2018-05-16 11:27:46 -04:00
Jarrod Johnson f1e29393df Succeed in pushing config to followers from leader
Still more work to be done for multiple transactions.
2018-05-16 11:27:46 -04:00
Jarrod Johnson 81bb16476c Apply changes from leader subscription 2018-05-16 11:27:46 -04:00
Jarrod Johnson 035f10e7d7 Rough draft for ongoing syncronization
Putting thoughts down on how xmit will work, will add recv and relay,
do some testing, and then decide how much can be done to apply it neatly
to the various points.
2018-05-16 11:27:46 -04:00
Jarrod Johnson 830e6bb4e4 Clear configuration prior to sync 2018-05-16 11:27:46 -04:00
Jarrod Johnson 1eb542f6a8 Actually execute replicate-on-connect
This creates a duplicate of the leader.
2018-05-16 11:27:46 -04:00
Jarrod Johnson b733049a0c Add self to collective database
Database would omit initial leader otherwise.
2018-05-16 11:27:46 -04:00
Jarrod Johnson a4d80e4e3a Fixes to the connect draft
Needed to track it's own name, skip the banner and auth message...
2018-05-16 11:27:46 -04:00
Jarrod Johnson a69b5fbb50 For enroll, track the remote cert special
For reconnect, we will have collective objects.  However at
enroll time, we need to be special.
2018-05-16 11:27:46 -04:00
Jarrod Johnson 789dfe94d0 Fix missing import of eventlet
Unable to spawn the connect thread due to missing import.
2018-05-16 11:27:46 -04:00
Jarrod Johnson 73a376fd74 Fix backup of globals
Globals failed to open the backup file as writable, causing failure if
a global had been set.
2018-05-16 11:27:46 -04:00
Jarrod Johnson 57f390fd0a Draft for starting the databse replication
Does not actually heed the data, or implement ongoing relay of data back and forth.
2018-05-16 11:27:46 -04:00
Jarrod Johnson 641bc7344a Add hooks for collective mode and refactor
In support of config replication, need configmanager to do a few things
2018-05-16 11:27:46 -04:00
Jarrod Johnson af940c972f Add function to check address equivalence
As we start needing to compare addresses, provide a central function
to handle the various oddities associated with that.
2018-05-16 11:27:46 -04:00
Jarrod Johnson 468f00cd1c Rename 'joinchallenge' to 'enroll'
Seems like a better word to use.
2018-05-16 11:27:46 -04:00
Jarrod Johnson ce635068aa Begin the 'connect' collective operation
First check if we are current leader, reject if not, then if cert
is invalid, reject, then comes the TODO.
2018-05-16 11:27:46 -04:00
Jarrod Johnson 348c6a7c38 Add collective info to DB backup
Now persisted to disk *and* accessible to backup.
2018-05-16 11:27:46 -04:00
Jarrod Johnson a41a42ffd0 Persint collective info to disk
Additionally, simplify the concluding steps of the join conversation.
2018-05-16 11:27:46 -04:00
Jarrod Johnson 3a354a6300 Ensure the invitation works out to even multiple of 3 bytes
It's cosmetic, but a nice way to avoid '=' in the tokens.
2018-05-16 11:27:46 -04:00
Jarrod Johnson fb9925bb84 Fix encoding of the response proof
The response was not decoded, causing it to always fail.
2018-05-16 11:27:46 -04:00
Jarrod Johnson 78bf1d5acd Provide more feedback and fix some flow issues 2018-05-16 11:27:46 -04:00
Jarrod Johnson 8165b645d9 Fix invite process and unicode
Unicode strings do not fit with our world view, make them bytes.
2018-05-16 11:27:46 -04:00
Jarrod Johnson c2783b6734 Rename swarm to collective in setup.py.tmpl 2018-05-16 11:27:46 -04:00
Jarrod Johnson 033d59b04a Afetr some feedback, rename it 'collective' 2018-05-16 11:27:46 -04:00
Jarrod Johnson 4155954d1c Add swarm to setup.py
Make sure the swarm content is actually installed.
2018-05-16 11:27:46 -04:00
Jarrod Johnson 1b912c4365 Further advance the swarm concept
This marks the start of attempting to connect the invitation
to sockets and using the invitation to measure the certificates as
well as proving client knowledge of an invitation token.
2018-05-16 11:27:46 -04:00
Jarrod Johnson 5f9ee3d3c5 Migrate 'multimanager' to 'swarm'
It's easier to say 'swarm' and conveys the sense without confusion
of 'cluster' mode.
2018-05-16 11:27:46 -04:00
Jarrod Johnson cc9becea3b Add ability to get client certificates
Unfortunately, to pull off the target user experience, we
must register a custom client certificate validation to allow
us to not require a CA.
2018-05-16 11:27:46 -04:00
Jarrod Johnson 1e0cf7e9fb Create invitation management module
This facilitates the generation of invitations and logistics of proving
knowledge of the invitation and the integrity of the certificates.
peercert is to be gotten through getpeercert(binary_form=True) and
local cert through the util function to load from file, since we don't
have another way of getting local certificate.
2018-05-16 11:26:36 -04:00
Jarrod Johnson 7ebe9da24b Add utility function to get certificate from file
This can be used to get our own certificate, for use in the
multimanager membership establishment.
2018-05-16 11:26:36 -04:00
Jarrod Johnson 6cba560f6a Fix nodeconfig handling of general errors
nodeconfig was not handling errors in results well, fix this by
refactoring the nodefirmware facility into it.
2018-05-16 11:21:26 -04:00
Jarrod Johnson 08dcab4c72 Fix load of txcount
Mistake caused txcount not to restore from disk.
2018-05-14 16:25:27 -04:00
Jarrod Johnson 4f73ddc41e Start setting the stage for leader change on restart
Have connect() have a way to recover if leader is dead.

Also these will be involved in configmanager detected loss of leader
2018-05-14 16:22:27 -04:00
Jarrod Johnson 7a912b31cb Fix transaction count in collective
Slave members were not persisting the value to disk
2018-05-14 15:35:44 -04:00
Jarrod Johnson 297513bba7 Correct scope of currentleader 2018-05-11 15:03:41 -04:00
Jarrod Johnson aa2be98dc3 Fix not having currentleader set
A slave node now recognizes itself as such.
2018-05-11 15:00:18 -04:00
Jarrod Johnson 3cc9ee1a17 Block some early startup problems in collective 2018-05-11 14:49:46 -04:00
Jarrod Johnson 173f1eaf7e Include absent devices in the json of nodeinventory 2018-05-11 14:24:47 -04:00
Jarrod Johnson 0d4b1a4213 Amend json output
Have the nodeinventory json output in a bit more directly useful format,
rather than regarding the API structured JSON...
2018-05-11 14:18:04 -04:00
Jarrod Johnson 58155a47b5 Add --json to nodeinventory
Have nodeinventory have an option to output in json.
2018-05-11 13:49:54 -04:00
Jarrod Johnson 7fa431dbc9 Persist the transactioncount
Needed for eventually ascertaining the viability in selecting leader.
2018-05-11 11:53:45 -04:00
Jarrod Johnson e98ecd9867 Correct spelling error 2018-05-10 16:46:29 -04:00
Jarrod Johnson 5087c8bed5 Try to persist name as myname
hostname may not agree with the name chosen by user, in such a case
persist the name and use that, falling back to gethostname()
as needed.
2018-05-10 16:40:13 -04:00
Jarrod Johnson 0477ab7d85 Check and try to start collective on startup
Not yet good enough for a leader to rejoin, but enough for a follower
to rejoin automatically.
2018-05-10 16:19:46 -04:00
Jarrod Johnson 267d83e6e4 Have attrib set wait on all collective members
This will mean that it is reliable that a nodeattrib ; <command>
in delegation scenarios is guaranteed to execute in order.
2018-05-09 17:01:23 -04:00
Jarrod Johnson c962d10222 Allow slave collective drones to set
It works (once), but needs xid fix.
2018-05-08 16:53:59 -04:00
Jarrod Johnson e5f553801b Fix reuse of channel for receiving changes 2018-05-08 13:35:30 -04:00
Jarrod Johnson d11c716b6a Succeed in pushing config to followers from leader
Still more work to be done for multiple transactions.
2018-05-08 11:33:42 -04:00
Jarrod Johnson aec4e746e9 Apply changes from leader subscription 2018-05-04 16:12:53 -04:00
Jarrod Johnson a78aa6816c Rough draft for ongoing syncronization
Putting thoughts down on how xmit will work, will add recv and relay,
do some testing, and then decide how much can be done to apply it neatly
to the various points.
2018-05-04 15:16:30 -04:00
Jarrod Johnson 5abaddfe63 Clear configuration prior to sync 2018-05-04 12:19:45 -04:00
Jarrod Johnson ecfc56efde Actually execute replicate-on-connect
This creates a duplicate of the leader.
2018-05-03 14:03:56 -04:00
Jarrod Johnson ecfb4d68c5 Add self to collective database
Database would omit initial leader otherwise.
2018-05-03 13:27:48 -04:00
Jarrod Johnson 855241a043 Fixes to the connect draft
Needed to track it's own name, skip the banner and auth message...
2018-05-03 13:18:08 -04:00
Jarrod Johnson ce4c72eae2 For enroll, track the remote cert special
For reconnect, we will have collective objects.  However at
enroll time, we need to be special.
2018-05-03 11:18:02 -04:00
Jarrod Johnson c8961377ed Fix missing import of eventlet
Unable to spawn the connect thread due to missing import.
2018-05-01 16:51:57 -04:00
Jarrod Johnson 1b17c42cae Fix backup of globals
Globals failed to open the backup file as writable, causing failure if
a global had been set.
2018-05-01 16:18:31 -04:00
Jarrod Johnson 196e8d0d58 Draft for starting the databse replication
Does not actually heed the data, or implement ongoing relay of data back and forth.
2018-04-30 16:20:49 -04:00
Jarrod Johnson dbb50f0807 Add hooks for collective mode and refactor
In support of config replication, need configmanager to do a few things
2018-04-30 16:09:28 -04:00
Jarrod Johnson 1200f7b7a1 Add function to check address equivalence
As we start needing to compare addresses, provide a central function
to handle the various oddities associated with that.
2018-04-30 16:08:15 -04:00
Jarrod Johnson a2a0b5de2c Rename 'joinchallenge' to 'enroll'
Seems like a better word to use.
2018-04-30 11:37:23 -04:00
Jarrod Johnson c7b01e00b6 Begin the 'connect' collective operation
First check if we are current leader, reject if not, then if cert
is invalid, reject, then comes the TODO.
2018-04-26 16:34:54 -04:00
Jarrod Johnson 06fdc648b8 Add collective info to DB backup
Now persisted to disk *and* accessible to backup.
2018-04-26 16:03:03 -04:00
Jarrod Johnson de89803b9c Persint collective info to disk
Additionally, simplify the concluding steps of the join conversation.
2018-04-26 15:48:15 -04:00
Jarrod Johnson d38d9204a7 Ensure the invitation works out to even multiple of 3 bytes
It's cosmetic, but a nice way to avoid '=' in the tokens.
2018-04-26 11:35:06 -04:00
Jarrod Johnson 1deb44021f Fix encoding of the response proof
The response was not decoded, causing it to always fail.
2018-04-25 20:49:08 -04:00
Jarrod Johnson 619bbbca96 Provide more feedback and fix some flow issues 2018-04-25 16:47:42 -04:00
Jarrod Johnson 8246ebdd2b Fix invite process and unicode
Unicode strings do not fit with our world view, make them bytes.
2018-04-25 14:21:12 -04:00
Jarrod Johnson a94a724fe0 Rename swarm to collective in setup.py.tmpl 2018-04-25 13:31:26 -04:00
Jarrod Johnson 6b9aed3722 Afetr some feedback, rename it 'collective' 2018-04-25 13:01:59 -04:00
Jarrod Johnson a3de8b9374 Add swarm to setup.py
Make sure the swarm content is actually installed.
2018-04-24 15:58:40 -04:00
Jarrod Johnson c8e5808daf Further advance the swarm concept
This marks the start of attempting to connect the invitation
to sockets and using the invitation to measure the certificates as
well as proving client knowledge of an invitation token.
2018-04-24 14:44:45 -04:00
Jarrod Johnson f3a0ccbff8 Migrate 'multimanager' to 'swarm'
It's easier to say 'swarm' and conveys the sense without confusion
of 'cluster' mode.
2018-04-24 12:59:24 -04:00
Jarrod Johnson 27c1355a4f Add ability to get client certificates
Unfortunately, to pull off the target user experience, we
must register a custom client certificate validation to allow
us to not require a CA.
2018-04-24 11:00:23 -04:00
Jarrod Johnson 7909f9e003 Switch to explicit SSL context when possible
This allows more fine grained control over the security parameters of
the TLS connection.
2018-04-23 14:18:51 -04:00
Jarrod Johnson 97d38efb29 Merge branch 'master' into clustertime 2018-04-23 11:13:22 -04:00
Jarrod Johnson 14ff33a44a Only activate the remote API socket if user makes cert
This prevents the useless networking socket from being opened
when it cannot be used.  This means most implementations will not
have an extra port to explain unless the user goes through the work
and knows what it would be.
2018-04-20 19:30:15 -04:00
Jarrod Johnson 78bdac474f Merge branch 'master' into clustertime 2018-04-20 13:25:16 -04:00
Jarrod Johnson 0481f7889b Make macmap api case insensitive
This helps usability of the api.
2018-04-20 13:25:02 -04:00
Jarrod Johnson 123a1a9dc1 Merge branch 'clustertime' of github.com:jjohnson42/confluent into clustertime 2018-04-17 11:10:13 -04:00
Jarrod Johnson fa24622704 Merge branch 'master' into clustertime 2018-04-17 11:09:59 -04:00
Jarrod Johnson a1156097d2 Add facility to disable autosense
discovery autosense at scale may produce undesirable performance.
Provide an interface to turn off the autosense.

If autosense is off, manual scan can still be performed.
2018-04-13 16:54:27 -04:00
Jarrod Johnson af72d0e71a Update the discovery lookup tables on node add/remove
This will mitigate stale mappings in the discovery process.
2018-04-12 17:05:06 -04:00
Jarrod Johnson 008f8e22ae Abort traversing gap in SMM chain
Once there is a gap, the next hop in the chain will be ambiguous.
Discovery must always precede from the front-most chassis.
2018-04-12 15:45:07 -04:00
Jarrod Johnson 39ee0da879 Fix makesetup for confluent_client
Fixing the redundant __init__.py led to no __init__.py, fix
that mistake.
2018-04-10 16:11:14 -04:00
Jarrod Johnson fc7b26eaf7 Remove __init__.py from tracking in client 2018-04-10 16:09:26 -04:00
Jarrod Johnson 91238f1dcb Clean up pure python packaging
Fix __init__.py redundancy, update requirements to current state
of affairs.
2018-04-10 16:06:37 -04:00
Jarrod Johnson 7f29d3a48f Merge branch 'master' into clustertime 2018-04-10 15:11:59 -04:00
Jarrod Johnson 76a4a91351 Fix pyparsing rpm name
Accept another likely formulation of an rpm name for
the package.
2018-04-10 15:11:20 -04:00
Jarrod Johnson 5ca52ff03b Handle interruptions to select such as resize
Resize can cause an interrupted operation on stdin, handle that.
2018-04-09 10:48:06 -04:00
Jarrod Johnson bd40f2f4a6 Fix mistake in indexing of url 2018-03-27 17:11:35 -04:00
Jarrod Johnson 66e8ce2dde Merge branch 'master' of github.com:jjohnson42/confluent 2018-03-27 16:35:31 -04:00
Jarrod Johnson 3dd86c71fd Add bmc.hostname to nodeconfig 2018-03-27 16:32:37 -04:00
Jarrod Johnson f97c39cea4 Add hostname to api
The hostname of the BMC is added to the api.
2018-03-27 15:51:14 -04:00
Jarrod Johnson 6671b9aad3 Provide cleaner behavior on timeouts
If a timeout occurred outside of a keeplaive, provide
a more consistent message about the situation.
2018-03-23 08:27:27 -04:00
Jarrod Johnson f88e0bca4c Fix nodeshell hang on incomplete lines
readline would hang because the filehandle was really not ready.
2018-03-19 08:45:13 -04:00
Jarrod Johnson afd366f134 Create invitation management module
This facilitates the generation of invitations and logistics of proving
knowledge of the invitation and the integrity of the certificates.
peercert is to be gotten through getpeercert(binary_form=True) and
local cert through the util function to load from file, since we don't
have another way of getting local certificate.
2018-03-15 19:22:03 -04:00
Jarrod Johnson ad0a7de1e3 Add utility function to get certificate from file
This can be used to get our own certificate, for use in the
multimanager membership establishment.
2018-03-15 18:59:49 -04:00
Jarrod Johnson 026a027603 Fix normalizing unicode in dicts with lists
If there's a list in a list, normalize that as well.
2018-03-15 12:55:32 -04:00
Jarrod Johnson 308db99dbb Fix inconsistent dict member extension
If two portions of a list come back piecewise from the plugin that
are both lists, extend them rather than making a nested list.
2018-03-15 12:09:45 -04:00
Jarrod Johnson a20b0abb43 Do not clear the buffer on superfluous reopen
If someone does a reopen, try to preserve the buffer, unless connect
proves there to be a deeper issue.  The risk of staleness is low, but
the experience of the whole screen clearing is tricky.  This was not
such an issue at the time, but using pyte causes clearbuffer to also
clear connected client terminals.
2018-03-14 17:00:44 -04:00
Jarrod Johnson 7413c44df8 Fix manual discovery
In manual discovery, maccount is not a field in the info, as no macmap
processing is done in manual.
2018-03-14 09:27:29 -04:00
Jarrod Johnson 463f61fac7 Modify XSS-Protection directive 2018-03-12 13:41:18 -04:00
Jarrod Johnson 0f60fc6df7 Fix uninitialized self._prevdict
self._prevdict was referenced without initialization.
2018-03-07 10:21:35 -05:00
Jarrod Johnson 110820e7b7 Revert "Accommodate XCC firmware behavior"
This reverts commit 9baa1f5652.
2018-03-06 15:52:53 -05:00
Jarrod Johnson 71214eb613 Revert "Correct indentation"
This reverts commit a2163244db.
2018-03-06 15:52:45 -05:00
Jarrod Johnson 889eda3d96 Merge remote-tracking branch 'upstream/master' 2018-03-06 11:25:54 -05:00
Jarrod Johnson 7593d21a87 Add missing exceptions import
exc was not imported
2018-03-06 11:25:36 -05:00
Jarrod Johnson 972801d41f Merge pull request #95 from aduffy19/nodepowerUpdate
Add previous option to nodepower command
2018-03-05 15:44:28 -05:00
Amanda Duffy b49531dfa5 Add previous option to nodepower command 2018-03-05 15:41:28 -05:00
53 changed files with 2411 additions and 309 deletions
+20 -3
View File
@@ -21,6 +21,7 @@
import optparse
import os
import select
import sys
path = os.path.dirname(os.path.realpath(__file__))
@@ -70,6 +71,9 @@ def print_current():
sys.stdout.flush()
fullline = sys.stdin.readline()
printpending = True
clearpending = False
holdoff = 0
while fullline:
for line in fullline.split('\n'):
if not line:
@@ -78,9 +82,22 @@ while fullline:
line = 'UNKNOWN: ' + line
grouped.add_line(*line.split(': ', 1))
if options.watch:
sys.stdout.write('\x1b[2J\x1b[;H') # clear screen
print_current()
if not holdoff:
holdoff = os.times()[4] + 0.250
if (holdoff < os.times()[4] or
not select.select((sys.stdin,), (), (), 0.250)[0]):
# print now, nothing pending
holdoff = 0
sys.stdout.write('\x1b[2J\x1b[;H') # clear screen
print_current()
printpending = False
clearpending = True
else:
printpending = True
fullline = sys.stdin.readline()
if not options.watch:
if printpending:
if clearpending:
sys.stdout.write('\x1b[2J\x1b[;H') # clear screen
print_current()
+13 -4
View File
@@ -763,7 +763,10 @@ def conserver_command(filehandle, localcommand):
def get_command_bytes(filehandle, localcommand, cmdlen):
while len(localcommand) < cmdlen:
ready, _, _ = select.select((filehandle,), (), (), 1)
try:
ready, _, _ = select.select((filehandle,), (), (), 1)
except select.error:
ready = ()
if ready:
localcommand += filehandle.read()
return localcommand
@@ -776,7 +779,10 @@ def check_escape_seq(currinput, filehandle):
sys.stdout.flush()
return conserver_command(
filehandle, currinput[len(conserversequence):])
ready, _, _ = select.select((filehandle,), (), (), 3)
try:
ready, _, _ = select.select((filehandle,), (), (), 3)
except select.error:
ready = ()
if not ready: # 3 seconds of no typing
break
currinput += filehandle.read()
@@ -866,8 +872,11 @@ def check_power_state():
while inconsole or not doexit:
if inconsole:
rdylist, _, _ = select.select(
(sys.stdin, session.connection), (), (), 10)
try:
rdylist, _, _ = select.select(
(sys.stdin, session.connection), (), (), 10)
except select.error:
rdylist = ()
for fh in rdylist:
if fh == session.connection:
# this only should get called in the
+1 -1
View File
@@ -35,7 +35,7 @@ if path.startswith('/opt'):
import confluent.client as client
argparser = optparse.OptionParser(
usage='''\n %prog [-b] noderange [list of attributes] \
usage='''\n %prog [-b] noderange [list of attributes or 'all'] \
\n %prog -c noderange <list of attributes> \
\n %prog -e noderange <attribute names to set> \
\n %prog noderange attribute1=value1 attribute2=value,...
+11 -9
View File
@@ -67,6 +67,8 @@ cfgpaths = {
'bmc.ipv4_gateway': (
'configuration/management_controller/net_interfaces/management',
'ipv4_gateway'),
'bmc.hostname': (
'configuration/management_controller/hostname', 'hostname'),
}
autodeps = {
@@ -185,12 +187,10 @@ if setmode:
for path in updatebypath:
for fr in session.update('/noderange/{0}/{1}'.format(noderange, path),
updatebypath[path]):
for node in fr['databynode']:
rcode |= client.printerror(fr)
for node in fr.get('databynode', []):
r = fr['databynode'][node]
if 'error' in r:
sys.stderr.write(node + ': ' + r['error'] + '\n')
if 'errorcode' in r:
rcode |= r['errorcode']
rcode |= client.printerror(r, node)
if 'value' not in r:
continue
keyval = r['value']
@@ -202,12 +202,14 @@ else:
for path in queryparms:
if options.comparedefault:
continue
client.print_attrib_path(path, session, list(queryparms[path]),
NullOpt(), queryparms[path])
rc = client.print_attrib_path(path, session, list(queryparms[path]),
NullOpt(), queryparms[path])
if rc:
sys.exit(rc)
if printsys or options.exclude:
if printsys == 'all':
printsys = []
path = '/noderange/{0}/configuration/system/all'.format(noderange)
client.print_attrib_path(path, session, printsys,
options)
rcode = client.print_attrib_path(path, session, printsys,
options)
sys.exit(rcode)
+1 -1
View File
@@ -47,7 +47,7 @@ session = client.Command()
exitcode = 0
attribs = {'name': noderange}
for arg in args[1:]:
key, val = arg.split('=')
key, val = arg.split('=', 1)
attribs[key] = val
for r in session.create('/noderange/', attribs):
if 'error' in r:
+3 -14
View File
@@ -35,18 +35,6 @@ import confluent.screensqueeze as sq
exitcode = 0
def printerror(res, node=None):
global exitcode
if 'errorcode' in res:
exitcode = res['errorcode']
if 'error' in res:
if node:
sys.stderr.write('{0}: {1}\n'.format(node, res['error']))
else:
sys.stderr.write('{0}\n'.format(res['error']))
if 'errorcode' not in res:
exitcode = 1
def printfirm(node, prefix, data):
if 'model' in data:
@@ -139,16 +127,17 @@ def update_firmware(session, filename):
sys.stderr.write('{0}: {1}\n'.format(node, noderrs[node]))
def show_firmware(session):
global exitcode
firmware_shown = False
for component in components:
for res in session.read(
'/noderange/{0}/inventory/firmware/all/{1}'.format(
noderange, component)):
printerror(res)
exitcode |= client.printerror(res)
if 'databynode' not in res:
continue
for node in res['databynode']:
printerror(res['databynode'][node], node)
exitcode |= client.printerror(res['databynode'][node], node)
if 'firmware' not in res['databynode'][node]:
continue
for inv in res['databynode'][node]['firmware']:
+1 -1
View File
@@ -47,7 +47,7 @@ session = client.Command()
exitcode = 0
attribs = {'name': noderange}
for arg in args[1:]:
key, val = arg.split('=')
key, val = arg.split('=', 1)
attribs[key] = val
for r in session.create('/nodegroups/', attribs):
if 'error' in r:
+25 -3
View File
@@ -16,6 +16,7 @@
# limitations under the License.
import codecs
import json
import optparse
import os
import re
@@ -89,9 +90,10 @@ url = '/noderange/{0}/inventory/hardware/all/all'
argparser = optparse.OptionParser(
usage="Usage: %prog <noderange> [serial|model|uuid|mac]")
argparser.add_option('-j', '--json', action='store_true', help='Output JSON')
(options, args) = argparser.parse_args()
try:
noderange = sys.argv[1]
noderange = args[0]
except IndexError:
argparser.print_help()
sys.exit(1)
@@ -114,6 +116,8 @@ if len(args) > 1:
filters.append(re.compile('mac address'))
url = '/noderange/{0}/inventory/hardware/all/all'
try:
if options.json:
databynode = {}
session = client.Command()
for res in session.read(url.format(noderange)):
printerror(res)
@@ -127,7 +131,12 @@ try:
prefix = inv['name']
if not inv['present']:
if not filters:
print '{0}: {1}: Not Present'.format(node, prefix)
if options.json:
if node not in databynode:
databynode[node] = {}
databynode[node][prefix] = inv
else:
print '{0}: {1}: Not Present'.format(node, prefix)
continue
info = inv['information']
info.pop('board_extra', None)
@@ -136,6 +145,11 @@ try:
info.pop('product_extra', None)
if 'memory_type' in info:
if not filters:
if options.json:
if node not in databynode:
databynode[node] = {}
databynode[node][prefix] = inv
continue
print_mem_info(node, prefix, info)
continue
for datum in info:
@@ -147,9 +161,17 @@ try:
continue
if info[datum] is None:
continue
if options.json:
if node not in databynode:
databynode[node] = {}
databynode[node][prefix] = inv
break
print(u'{0}: {1} {2}: {3}'.format(node, prefix,
pretty(datum),
info[datum]))
if options.json:
print(json.dumps(databynode, sort_keys=True, indent=4,
separators=(',', ': ')))
except KeyboardInterrupt:
print('')
sys.exit(exitcode)
sys.exit(exitcode)
+22 -5
View File
@@ -34,6 +34,9 @@ import confluent.client as client
argparser = optparse.OptionParser(
usage="Usage: %prog [options] noderange "
"([status|on|off|shutdown|boot|reset])")
argparser.add_option('-p', '--showprevious', dest='previous',
action='store_true', default=False,
help='Show previous power state')
(options, args) = argparser.parse_args()
try:
noderange = args[0]
@@ -42,11 +45,11 @@ except IndexError:
sys.exit(1)
client.check_globbing(noderange)
setstate = None
if len(sys.argv) > 2:
if len(args) > 1:
if setstate == 'softoff':
setstate = 'shutdown'
elif not sys.argv[2] in ('stat', 'state', 'status'):
setstate = sys.argv[2]
elif not args[1] in ('stat', 'state', 'status'):
setstate = args[1]
if setstate not in (None, 'on', 'off', 'shutdown', 'boot', 'reset'):
argparser.print_help()
@@ -54,5 +57,19 @@ if setstate not in (None, 'on', 'off', 'shutdown', 'boot', 'reset'):
session = client.Command()
exitcode = 0
session.add_precede_key('oldstate')
sys.exit(
session.simple_noderange_command(noderange, '/power/state', setstate))
if options.previous:
# get previous states
prev = {}
for rsp in session.read("/noderange/{0}/power/state".format(noderange)):
# gets previous (current) states
databynode = rsp["databynode"]
for node in databynode:
prev[node] = databynode[node]["state"]["value"]
# add dictionary to session
session.add_precede_dict(prev)
sys.exit(session.simple_noderange_command(noderange, '/power/state', setstate))
+1 -1
View File
@@ -86,7 +86,7 @@ def run():
desc = pipedesc[r]
node = desc['node']
data = True
while data and select.select([r], [], [], 0):
while data and select.select([r], [], [], 0)[0]:
data = r.readline()
if data:
if desc['type'] == 'stdout':
+30 -3
View File
@@ -33,6 +33,21 @@ _attraliases = {
'bmcpass': 'secret.hardwaremanagementpassword',
}
def printerror(res, node=None):
exitcode = 0
if 'errorcode' in res:
exitcode = res['errorcode']
if 'error' in res:
if node:
sys.stderr.write('{0}: {1}\n'.format(node, res['error']))
else:
sys.stderr.write('{0}\n'.format(res['error']))
if 'errorcode' not in res:
exitcode = 1
return exitcode
def cprint(txt):
print(txt)
sys.stdout.flush()
@@ -53,6 +68,7 @@ def _parseserver(string):
class Command(object):
def __init__(self, server=None):
self._prevdict = None
self._prevkeyname = None
self.connection = None
self._currnoderange = None
@@ -88,6 +104,9 @@ class Command(object):
def add_precede_key(self, keyname):
self._prevkeyname = keyname
def add_precede_dict(self, dict):
self._prevdict = dict
def handle_results(self, ikey, rc, res, errnodes=None):
if 'error' in res:
if errnodes is not None:
@@ -120,6 +139,9 @@ class Command(object):
if self._prevkeyname and self._prevkeyname in res[node]:
cprint('{0}: {2}->{1}'.format(
node, val, res[node][self._prevkeyname]['value']))
elif self._prevdict and node in self._prevdict:
cprint('{0}: {2}->{1}'.format(
node, val, self._prevdict[node]))
else:
cprint('{0}: {1}'.format(node, val))
return rc
@@ -239,8 +261,7 @@ class Command(object):
certreqs = ssl.CERT_NONE
knownhosts = True
self.connection = ssl.wrap_socket(self.connection, ca_certs=cacert,
cert_reqs=certreqs,
ssl_version=ssl.PROTOCOL_TLSv1)
cert_reqs=certreqs)
if knownhosts:
certdata = self.connection.getpeercert(binary_form=True)
fingerprint = 'sha512$' + hashlib.sha512(certdata).hexdigest()
@@ -316,6 +337,12 @@ def print_attrib_path(path, session, requestargs, options, rename=None):
for attr, val in sorted(
res['databynode'][node].items(),
key=lambda (k, v): v.get('sortid', k) if isinstance(v, dict) else k):
if attr == 'error':
sys.stderr.write('{0}: Error: {1}\n'.format(node, val))
continue
if attr == 'errorcode':
exitcode |= val
continue
seenattributes.add(attr)
if rename:
printattr = rename.get(attr, attr)
@@ -503,7 +530,7 @@ def updateattrib(session, updateargs, nodetype, noderange, options):
if "=" in updateargs[1]:
try:
for val in updateargs[1:]:
val = val.split('=')
val = val.split('=', 1)
if val[0][-1] in (',', '-', '^'):
key = val[0][:-1]
if val[0][-1] == ',':
+30 -6
View File
@@ -20,6 +20,11 @@ from datetime import datetime
import json
import struct
try:
unicode
except NameError:
unicode = str
def decodestr(value):
ret = None
try:
@@ -38,17 +43,28 @@ def unicode_dictvalues(dictdata):
elif isinstance(dictdata[key], datetime):
dictdata[key] = dictdata[key].strftime('%Y-%m-%dT%H:%M:%S')
elif isinstance(dictdata[key], list):
for i in xrange(len(dictdata[key])):
if isinstance(dictdata[key][i], str):
dictdata[key][i] = decodestr(dictdata[key][i])
elif isinstance(dictdata[key][i], dict):
unicode_dictvalues(dictdata[key][i])
_unicode_list(dictdata[key])
elif isinstance(dictdata[key], dict):
unicode_dictvalues(dictdata[key])
def _unicode_list(currlist):
for i in xrange(len(currlist)):
if isinstance(currlist[i], str):
currlist[i] = decodestr(currlist[i])
elif isinstance(currlist[i], dict):
unicode_dictvalues(currlist[i])
elif isinstance(currlist[i], list):
_unicode_list(currlist[i])
def send(handle, data):
if isinstance(data, str):
if isinstance(data, unicode):
try:
data = data.encode('utf-8')
except AttributeError:
pass
if isinstance(data, str) or isinstance(data, unicode):
# plain text, e.g. console data
tl = len(data)
if tl == 0:
@@ -74,6 +90,14 @@ def send(handle, data):
handle.sendall(struct.pack("!I", tl))
handle.sendall(sdata)
def recvall(handle, size):
rd = handle.recv(size)
while len(rd) < size:
nd = handle.recv(size - len(rd))
if not nd:
raise Exception("Error reading data")
rd += nd
return rd
def recv(handle):
tl = handle.recv(4)
+7 -3
View File
@@ -3,7 +3,7 @@ nodeattrib(8) -- List or change confluent nodes attributes
## SYNOPSIS
`nodeattrib [-b] <noderange> [<nodeattribute>...]`
`nodeattrib [-b] <noderange> [all|<nodeattribute>...]`
`nodeattrib <noderange> [<nodeattribute1=value1> <nodeattribute2=value2> ...]`
`nodeattrib -c <noderange> <nodeattribute1> <nodeattribute2> ...`
`nodeattrib -e <noderange> <nodeattribute1> <nodeattribute2> ...`
@@ -19,11 +19,15 @@ displayed. If `-b` is specified, it will also display information on
how inherited and expression based attributes are defined. Attributes can be
straightforward values, or an expression as documented in nodeattribexpressions(5).
For a full list of attributes, run `nodeattrib <node> all` against a node.
If `-c` is specified, this will set the nodeattribute to a null valid.
If `-c` is specified, this will set the nodeattribute to a null value.
This is different from setting the value to an empty string.
If the word all is specified, then all available attributes are given.
Omitting any attribute name or the word 'all' will display only attributes
that are currently set.
For the `groups` attribute, it is possible to add a group by doing
`groups,=<newgroup>`` and to remove by doing `groups^=<oldgroup>`
`groups,=<newgroup>` and to remove by doing `groups^=<oldgroup>`
Note that `nodeattrib <group>` will likely not provide the expected behavior.
See nodegroupattrib(8) command on how to manage attributes on a group level.
+2 -1
View File
@@ -15,7 +15,8 @@ are mounted in an insecure fashion. http is insecure, and https is also
insecure when no meaningful certificate validation is performed. Currently
there is no action that can change this, and this is purely informational. A
future version of software may provide a means to increase security of attached
remote media.
remote media. If no media is mounted, this will provide no output, error
conditions will result in output to standard error.
`detachall` removes all the currently provided media to the host. This unlinks
remote media from urls and deletes uploaded media from the BMC.
+6 -1
View File
@@ -4,7 +4,7 @@ nodepower(8) -- Check or change power state of confluent nodes
## SYNOPSIS
`nodepower <noderange>`
`nodepower <noderange> [on|off|boot|shutdown|reset|status]`
`nodepower <noderange> [-p] [on|off|boot|shutdown|reset|status]`
## DESCRIPTION
@@ -26,6 +26,11 @@ respond.
off will not react to this request.
* `status`: Behave identically to having no argument passed at all.
## OPTIONS
* `-p`, '--showprevious':
Show previous power state for all directives that may change power state.
## EXAMPLES
* Get power state of nodes n1 through n4:
`# nodepower n1-n4`
+4
View File
@@ -48,6 +48,10 @@ themselves, see nodeshell(8).
`n4: 01 10 00`
`n2: 01 10 00`
* If wanting to use literal {} in the command, they must be escaped by doubling:
`# noderun n1-n4 "echo {node} | awk '{{print $1}}'"`
## SEE ALSO
nodeshell(8)
+4 -1
View File
@@ -27,7 +27,10 @@ as stderr, unlike psh which combines all stdout and stderr into stdout.
`n4: hi`
* Setting a new static ip address temporarily on secondary interface of four nodes:
`# nodeshell n1-n4 ifconfig eth1 172.30.93.{n1}`
`# nodeshell n1-n4 ifconfig eth1 172.30.93.{n1}`
* If wanting to use literal {} in the command, they must be escaped by doubling:
`# nodeshell n1-n4 "ps | awk '{{print $1}}'"`
## SEE ALSO
@@ -0,0 +1,35 @@
#!/usr/bin/env python
# This is a sample python script for going through all observed mac addresses
# and assuming they are BMC related and printing nodeattrib commands
# for each node to access the bmc using the interface specified on the command
# line
# Not necessarily as useful if there may be mistakes in the
# net.switch/net.switchport attributes, but a handy utility in a pinch when
# you really know
import confluent.client as cl
import socket
import struct
c = cl.Command()
macs = []
interface = sys.argv[1]
for mac in c.read('/networking/macs/by-mac/'):
macs.append(mac['item']['href'])
for mac in macs:
macinfo = list(c.read('/networking/macs/by-mac/{0}'.format(mac)))[0]
if 'possiblenode' in macinfo and macinfo['possiblenode']:
if macinfo['macsonport'] > 1:
print('#Ambiguous set of macs on port for ' + macinfo[
'possiblenode'])
prefix = int(mac.replace('-', '')[:6], 16) ^ 0b100000000000000000
prefix = prefix << 8
prefix |= 0xff
suffix = int(mac.replace('-', '')[6:], 16)
suffix |= 0xfe000000
rawn = struct.pack('!QLL', 0xfe80000000000000, prefix, suffix)
bmc = socket.inet_ntop(socket.AF_INET6, rawn)
print('nodeattrib {0} bmc={1}%{2}'.format(macinfo['possiblenode'],
bmc, interface))
+121
View File
@@ -0,0 +1,121 @@
#!/usr/bin/env python
import argparse
import errno
import os
import socket
import subprocess
import sys
path = os.path.dirname(os.path.realpath(__file__))
path = os.path.realpath(os.path.join(path, '..', 'lib', 'python'))
if path.startswith('/opt'):
# if installed into system path, do not muck with things
sys.path.append(path)
import confluent.client as client
import confluent.tlvdata as tlvdata
try:
input = raw_input
except NameError:
pass
def make_certificate():
umask = os.umask(0077)
try:
os.makedirs('/etc/confluent/cfg')
except OSError as e:
if e.errno == errno.EEXIST and os.path.isdir('/etc/confluent/cfg'):
pass
else:
raise
if subprocess.check_call(
'openssl ecparam -name secp384r1 -genkey -out '
'/etc/confluent/privkey.pem'.split(' ')):
raise Exception('Error generating private key')
if subprocess.check_call('openssl req -new -x509 -key '
'/etc/confluent/privkey.pem -days 7300 -out '
'/etc/confluent/srvcert.pem -subj /CN='
'{0}'.format(socket.gethostname()).split(' ')):
raise Exception('Error generating certificate')
print('Certificate generated successfully')
os.umask(umask)
def show_invitation(name):
if not os.path.exists('/etc/confluent/srvcert.pem'):
make_certificate()
s = client.Command().connection
tlvdata.send(s, {'collective': {'operation': 'invite', 'name': name}})
invite = tlvdata.recv(s)['collective']
if 'error' in invite:
sys.stderr.write(invite['error'] + '\n')
return
print('{0}'.format(invite['invitation']))
def join_collective(server, invitation):
if not os.path.exists('/etc/confluent/srvcert.pem'):
make_certificate()
s = client.Command().connection
while not invitation:
invitation = raw_input('Paste the invitation here: ')
tlvdata.send(s, {'collective': {'operation': 'join',
'invitation': invitation,
'server': server}})
res = tlvdata.recv(s)
print(res.get('collective',
{'status': 'Unknown response: ' + repr(res)})['status'])
def show_collective():
s = client.Command().connection
tlvdata.send(s, {'collective': {'operation': 'show'}})
res = tlvdata.recv(s)
if 'error' in res['collective']:
print(res['collective']['error'])
return
if 'quorum' in res['collective']:
print('Quorum: {0}'.format(res['collective']['quorum']))
print('Leader: {0}'.format(res['collective']['leader']))
if 'active' in res['collective']:
if res['collective']['active']:
print('Active collective members:')
for member in res['collective']['active']:
print(' {0}'.format(member))
if res['collective']['offline']:
print('Offline collective members:')
for member in res['collective']['offline']:
print(' {0}'.format(member))
else:
print('Run collective show on leader for more data')
def main():
a = argparse.ArgumentParser(description='Confluent server utility')
sp = a.add_subparsers(dest='command')
gc = sp.add_parser('gencert', help='Generate Confluent Certificates for '
'collective mode and remote CLI access')
sl = sp.add_parser('show', help='Show information about the collective')
ic = sp.add_parser('invite', help='Generate a invitation to allow a new '
'confluent instance to join as a '
'collective member')
ic.add_argument('name', help='Name of server to invite to join the '
'collective')
jc = sp.add_parser('join', help='Join a collective')
jc.add_argument('server', help='A server currently in the collective')
jc.add_argument('-i', help='Invitation provided by runniing invite on an '
'existing collective member')
cmdset = a.parse_args()
if cmdset.command == 'gencert':
make_certificate()
elif cmdset.command == 'invite':
show_invitation(cmdset.name)
elif cmdset.command == 'join':
join_collective(cmdset.server, cmdset.i)
elif cmdset.command == 'show':
show_collective()
if __name__ == '__main__':
main()
+30 -4
View File
@@ -16,6 +16,7 @@
# limitations under the License.
import getpass
import optparse
import sys
import os
@@ -32,6 +33,8 @@ argparser = optparse.OptionParser(
usage="Usage: %prog [options] [dump|restore] [path]")
argparser.add_option('-p', '--password',
help='Password to use to protect/unlock a protected dump')
argparser.add_option('-i', '--interactivepassword', help='Prompt for password',
action='store_true')
argparser.add_option('-r', '--redact', action='store_true',
help='Redact potentially sensitive data rather than store')
argparser.add_option('-u', '--unprotected', action='store_true',
@@ -39,6 +42,14 @@ argparser.add_option('-u', '--unprotected', action='store_true',
' the key information. Fields will be encrypted, '
'but keys.json will contain unencrypted decryption'
' keys that may be used to read the dump')
argparser.add_option('-s', '--skipkeys', action='store_true',
help='This specifies to dump the encrypted data without '
'dumping the keys needed to decrypt it. This is '
'suitable for an automated incremental backup, '
'where an earlier password protected dump has a '
'protected keys.json file, and only the protected '
'data is needed. keys do not change and as such '
'they do not require incremental backup')
(options, args) = argparser.parse_args()
if len(args) != 2 or args[0] not in ('dump', 'restore'):
argparser.print_help()
@@ -51,17 +62,32 @@ if args[0] == 'restore':
if pid is not None:
print("Confluent is running, must shut down to restore db")
sys.exit(1)
cfm.restore_db_from_directory(dumpdir, options.password)
try:
cfm.restore_db_from_directory(dumpdir, options.password)
except Exception as e:
print(str(e))
sys.exit(1)
elif args[0] == 'dump':
if options.password is None and not (options.unprotected or options.redact):
password = options.password
if not password and options.interactivepassword:
passcfm = None
while passcfm is None or password != passcfm:
password = getpass.getpass(
'Enter password to protect the backup: ')
passcfm = getpass.getpass('Confirm password to protect the backup: ')
if password is None and not (options.unprotected or options.redact
or options.skipkeys):
print("Must indicate a password to protect or -u to opt opt of "
"secure value protection or -r to skip all protected data")
"secure value protection or -r to redact sensitive information, "
"or -s to do encrypted backup that requires keys.json from "
"another backup to restore.")
sys.exit(1)
os.umask(077)
main._initsecurity(conf.get_config())
if not os.path.exists(dumpdir):
os.makedirs(dumpdir)
cfm.dump_db_to_directory(dumpdir, options.password, options.redact)
cfm.dump_db_to_directory(dumpdir, options.password, options.redact,
options.skipkeys)
+4
View File
@@ -0,0 +1,4 @@
#!/bin/bash
umask 0077
openssl ecparam -name secp384r1 -genkey -out /etc/confluent/privkey.pem
openssl req -new -x509 -key /etc/confluent/privkey.pem -days 760 -out /etc/confluent/srvcert.pem -subj /CN=$(hostname)
-1
View File
@@ -1 +0,0 @@
__version__ = "1.5.0.dev395.ggacb3a20"
+1 -1
View File
@@ -22,7 +22,7 @@
import confluent.config.configmanager as configmanager
import eventlet
import eventlet.tpool
import Crypto.Protocol.KDF as KDF
import Cryptodome.Protocol.KDF as KDF
import hashlib
import hmac
import multiprocessing
@@ -0,0 +1,60 @@
# vim: tabstop=4 shiftwidth=4 softtabstop=4
# Copyright 2018 Lenovo
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
# This handles the process of generating and tracking/validating invites
import base64
import hashlib
import hmac
import os
pending_invites = {}
def create_server_invitation(servername):
servername = servername.encode('utf-8')
randbytes = (3 - ((len(servername) + 2) % 3)) % 3 + 64
invitation = os.urandom(randbytes)
pending_invites[servername] = invitation
return base64.b64encode(servername + b'@' + invitation)
def create_client_proof(invitation, mycert, peercert):
return hmac.new(invitation, peercert + mycert, hashlib.sha256).digest()
def check_server_proof(invitation, mycert, peercert, proof):
validproof = hmac.new(invitation, mycert + peercert, hashlib.sha256
).digest()
return proof == validproof
def check_client_proof(servername, mycert, peercert, proof):
servername = servername.encode('utf-8')
if servername not in pending_invites:
return False
invitation = pending_invites[servername]
validproof = hmac.new(invitation, mycert + peercert, hashlib.sha256
).digest()
if proof == validproof:
# We know that the client knew the secret, and that it measured our
# certificate, and thus calling code can bless the certificate, and
# we can forget the invitation
del pending_invites[servername]
# We now want to prove to the client that we also know the secret,
# and that we measured their certificate well
# Now to generate an answer...., reverse the cert order so our answer
# is different, but still proving things
return hmac.new(invitation, peercert + mycert, hashlib.sha256
).digest()
# The given proof did not verify the invitation
return False
@@ -0,0 +1,494 @@
# vim: tabstop=4 shiftwidth=4 softtabstop=4
# Copyright 2018 Lenovo
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
import base64
import confluent.collective.invites as invites
import confluent.config.configmanager as cfm
import confluent.exceptions as exc
import confluent.tlvdata as tlvdata
import confluent.util as util
import eventlet
import eventlet.green.socket as socket
import eventlet.green.ssl as ssl
import eventlet.green.threading as threading
import random
try:
import OpenSSL.crypto as crypto
except ImportError:
# while not always required, we use pyopenssl required for at least
# collective
crypto = None
currentleader = None
cfginitlock = None
follower = None
retrythread = None
class ContextBool(object):
def __init__(self):
self.active = False
def __enter__(self):
self.active = True
def __exit__(self, exc_type, exc_val, exc_tb):
self.active = False
connecting = ContextBool()
leader_init = ContextBool()
def connect_to_leader(cert=None, name=None, leader=None):
global currentleader
global cfginitlock
global follower
if cfginitlock is None:
cfginitlock = threading.RLock()
if leader is None:
leader = currentleader
try:
remote = connect_to_collective(cert, leader)
except socket.error:
return False
with connecting:
with cfginitlock:
tlvdata.recv(remote) # the banner
tlvdata.recv(remote) # authpassed... 0..
if name is None:
name = get_myname()
tlvdata.send(remote, {'collective': {'operation': 'connect',
'name': name,
'txcount': cfm._txcount}})
keydata = tlvdata.recv(remote)
if not keydata:
return False
if 'error' in keydata:
if 'backoff' in keydata:
eventlet.spawn_after(random.random(), connect_to_leader,
cert, name, leader)
return True
if 'leader' in keydata:
ldrc = cfm.get_collective_member_by_address(
keydata['leader'])
if ldrc and ldrc['name'] == name:
raise Exception("Redirected to self")
return connect_to_leader(name=name,
leader=keydata['leader'])
if 'txcount' in keydata:
return become_leader(remote)
print(keydata['error'])
return False
if follower is not None:
follower.kill()
cfm.stop_following()
follower = None
colldata = tlvdata.recv(remote)
globaldata = tlvdata.recv(remote)
dbi = tlvdata.recv(remote)
dbsize = dbi['dbsize']
dbjson = ''
while (len(dbjson) < dbsize):
ndata = remote.recv(dbsize - len(dbjson))
if not ndata:
try:
remote.close()
except Exception:
pass
raise Exception("Error doing initial DB transfer")
dbjson += ndata
cfm.clear_configuration()
try:
cfm._restore_keys(keydata, None, sync=False)
for c in colldata:
cfm._true_add_collective_member(c, colldata[c]['address'],
colldata[c]['fingerprint'],
sync=False)
for globvar in globaldata:
cfm.set_global(globvar, globaldata[globvar], False)
cfm._txcount = dbi.get('txcount', 0)
cfm.ConfigManager(tenant=None)._load_from_json(dbjson,
sync=False)
cfm.commit_clear()
except Exception:
cfm.stop_following()
cfm.rollback_clear()
raise
currentleader = leader
#spawn this as a thread...
follower = eventlet.spawn(follow_leader, remote)
return True
def follow_leader(remote):
global currentleader
cfm.follow_channel(remote)
# The leader has folded, time to startup again...
cfm.stop_following()
currentleader = None
eventlet.spawn_n(start_collective)
def connect_to_collective(cert, member):
remote = socket.create_connection((member, 13001))
# TLS cert validation is custom and will not pass normal CA vetting
# to override completely in the right place requires enormous effort, so just defer until after connect
remote = ssl.wrap_socket(remote, cert_reqs=ssl.CERT_NONE, keyfile='/etc/confluent/privkey.pem',
certfile='/etc/confluent/srvcert.pem')
if cert:
fprint = cert
else:
collent = cfm.get_collective_member_by_address(member)
fprint = collent['fingerprint']
if not util.cert_matches(fprint, remote.getpeercert(binary_form=True)):
# probably Janeway up to something
raise Exception("Certificate mismatch in the collective")
return remote
def get_myname():
try:
with open('/etc/confluent/cfg/myname', 'r') as f:
return f.read().strip()
except IOError:
myname = socket.gethostname()
with open('/etc/confluent/cfg/myname', 'w') as f:
f.write(myname)
return myname
def handle_connection(connection, cert, request, local=False):
global currentleader
global retrythread
operation = request['operation']
if cert:
cert = crypto.dump_certificate(crypto.FILETYPE_ASN1, cert)
else:
if not local:
return
if 'show' == operation:
if not list(cfm.list_collective()):
tlvdata.send(connection,
{'collective': {'error': 'Collective mode not '
'enabled on this '
'system'}})
return
if follower:
linfo = cfm.get_collective_member_by_address(currentleader)
remote = socket.create_connection((currentleader, 13001))
remote = ssl.wrap_socket(remote, cert_reqs=ssl.CERT_NONE,
keyfile='/etc/confluent/privkey.pem',
certfile='/etc/confluent/srvcert.pem')
cert = remote.getpeercert(binary_form=True)
if not (linfo and util.cert_matches(
linfo['fingerprint'],
cert)):
remote.close()
tlvdata.send(connection,
{'error': 'Invalid certificate, '
'redo invitation process'})
connection.close()
return
tlvdata.recv(remote) # ignore banner
tlvdata.recv(remote) # ignore authpassed: 0
tlvdata.send(remote,
{'collective': {'operation': 'getinfo',
'name': get_myname()}})
collinfo = tlvdata.recv(remote)
else:
collinfo = {}
populate_collinfo(collinfo)
try:
cfm.check_quorum()
collinfo['quorum'] = True
except exc.DegradedCollective:
collinfo['quorum'] = False
tlvdata.send(connection, {'collective': collinfo})
return
if 'invite' == operation:
try:
cfm.check_quorum()
except exc.DegradedCollective:
tlvdata.send(connection,
{'collective':
{'error': 'Collective does not have quorum'}})
return
#TODO(jjohnson2): Cannot do the invitation if not the head node, the certificate hand-carrying
#can't work in such a case.
name = request['name']
invitation = invites.create_server_invitation(name)
tlvdata.send(connection,
{'collective': {'invitation': invitation}})
connection.close()
if 'join' == operation:
invitation = request['invitation']
try:
invitation = base64.b64decode(invitation)
name, invitation = invitation.split('@', 1)
except Exception:
tlvdata.send(
connection,
{'collective':
{'status': 'Invalid token format'}})
connection.close()
return
host = request['server']
try:
remote = socket.create_connection((host, 13001))
# This isn't what it looks like. We do CERT_NONE to disable
# openssl verification, but then use the invitation as a
# shared secret to validate the certs as part of the join
# operation
remote = ssl.wrap_socket(remote, cert_reqs=ssl.CERT_NONE,
keyfile='/etc/confluent/privkey.pem',
certfile='/etc/confluent/srvcert.pem')
except Exception:
tlvdata.send(
connection,
{'collective':
{'status': 'Failed to connect to {0}'.format(host)}})
connection.close()
return
mycert = util.get_certificate_from_file(
'/etc/confluent/srvcert.pem')
cert = remote.getpeercert(binary_form=True)
proof = base64.b64encode(invites.create_client_proof(
invitation, mycert, cert))
tlvdata.recv(remote) # ignore banner
tlvdata.recv(remote) # ignore authpassed: 0
tlvdata.send(remote, {'collective': {'operation': 'enroll',
'name': name, 'hmac': proof}})
rsp = tlvdata.recv(remote)
if 'error' in rsp:
tlvdata.send(connection, {'collective':
{'status': rsp['error']}})
connection.close()
return
proof = rsp['collective']['approval']
proof = base64.b64decode(proof)
j = invites.check_server_proof(invitation, mycert, cert, proof)
if not j:
remote.close()
tlvdata.send(connection, {'collective':
{'status': 'Bad server token'}})
connection.close()
return
tlvdata.send(connection, {'collective': {'status': 'Success'}})
connection.close()
currentleader = rsp['collective']['leader']
f = open('/etc/confluent/cfg/myname', 'w')
f.write(name)
f.close()
eventlet.spawn_n(connect_to_leader, rsp['collective'][
'fingerprint'], name)
if 'enroll' == operation:
#TODO(jjohnson2): error appropriately when asked to enroll, but the master is elsewhere
mycert = util.get_certificate_from_file('/etc/confluent/srvcert.pem')
proof = base64.b64decode(request['hmac'])
myrsp = invites.check_client_proof(request['name'], mycert,
cert, proof)
if not myrsp:
tlvdata.send(connection, {'error': 'Invalid token'})
connection.close()
return
myrsp = base64.b64encode(myrsp)
fprint = util.get_fingerprint(cert)
myfprint = util.get_fingerprint(mycert)
cfm.add_collective_member(get_myname(),
connection.getsockname()[0], myfprint)
cfm.add_collective_member(request['name'],
connection.getpeername()[0], fprint)
myleader = get_leader(connection)
ldrfprint = cfm.get_collective_member_by_address(
myleader)['fingerprint']
tlvdata.send(connection,
{'collective': {'approval': myrsp,
'fingerprint': ldrfprint,
'leader': get_leader(connection)}})
if 'assimilate' == operation:
drone = request['name']
droneinfo = cfm.get_collective_member(drone)
if not util.cert_matches(droneinfo['fingerprint'], cert):
tlvdata.send(connection,
{'error': 'Invalid certificate, '
'redo invitation process'})
return
if request['txcount'] < cfm._txcount:
tlvdata.send(connection,
{'error': 'Refusing to be assimilated by inferior'
'transaction count',
'txcount': cfm._txcount})
return
eventlet.spawn_n(connect_to_leader, None, None,
leader=connection.getpeername()[0])
tlvdata.send(connection, {'status': 0})
connection.close()
if 'getinfo' == operation:
drone = request['name']
droneinfo = cfm.get_collective_member(drone)
if not (droneinfo and util.cert_matches(droneinfo['fingerprint'],
cert)):
tlvdata.send(connection,
{'error': 'Invalid certificate, '
'redo invitation process'})
connection.close()
return
collinfo = {}
populate_collinfo(collinfo)
tlvdata.send(connection, collinfo)
if 'connect' == operation:
drone = request['name']
droneinfo = cfm.get_collective_member(drone)
if not (droneinfo and util.cert_matches(droneinfo['fingerprint'],
cert)):
tlvdata.send(connection,
{'error': 'Invalid certificate, '
'redo invitation process'})
connection.close()
return
myself = connection.getsockname()[0]
if myself != get_leader(connection):
tlvdata.send(
connection,
{'error': 'Cannot assimilate, our leader is '
'in another castle', 'leader': currentleader})
connection.close()
return
if connecting.active:
tlvdata.send(connection, {'error': 'Connecting right now',
'backoff': True})
connection.close()
return
if request['txcount'] > cfm._txcount:
retire_as_leader()
tlvdata.send(connection,
{'error': 'Client has higher tranasaction count, '
'should assimilate me, connecting..',
'txcount': cfm._txcount})
eventlet.spawn_n(connect_to_leader, None, None,
connection.getpeername()[0])
connection.close()
return
if retrythread:
retrythread.cancel()
retrythread = None
with leader_init:
cfm.update_collective_address(request['name'],
connection.getpeername()[0])
tlvdata.send(connection, cfm._dump_keys(None, False))
tlvdata.send(connection, cfm._cfgstore['collective'])
tlvdata.send(connection, cfm.get_globals())
cfgdata = cfm.ConfigManager(None)._dump_to_json()
tlvdata.send(connection, {'txcount': cfm._txcount,
'dbsize': len(cfgdata)})
connection.sendall(cfgdata)
#tlvdata.send(connection, {'tenants': 0}) # skip the tenants for now,
# so far unused anyway
if not cfm.relay_slaved_requests(drone, connection):
if not retrythread: # start a recovery if everyone else seems
# to have disappeared
retrythread = eventlet.spawn_after(30 + random.random(),
start_collective)
# ok, we have a connecting member whose certificate checks out
# He needs to bootstrap his configuration and subscribe it to updates
def populate_collinfo(collinfo):
iam = get_myname()
collinfo['leader'] = iam
collinfo['active'] = list(cfm.cfgstreams)
activemembers = set(cfm.cfgstreams)
activemembers.add(iam)
collinfo['offline'] = []
for member in cfm.list_collective():
if member not in activemembers:
collinfo['offline'].append(member)
def try_assimilate(drone):
try:
remote = connect_to_collective(None, drone)
except socket.error:
# Oh well, unable to connect, hopefully the rest will be
# in order
return
tlvdata.send(remote, {'collective': {'operation': 'assimilate',
'name': get_myname(),
'txcount': cfm._txcount}})
tlvdata.recv(remote) # the banner
tlvdata.recv(remote) # authpassed... 0..
answer = tlvdata.recv(remote)
if answer and 'error' in answer:
connect_to_leader(None, None, leader=remote.getpeername()[0])
def get_leader(connection):
if currentleader is None or connection.getpeername()[0] == currentleader:
become_leader(connection)
return currentleader
def retire_as_leader():
global currentleader
cfm.stop_leading()
currentleader = None
def become_leader(connection):
global currentleader
global follower
if follower:
follower.kill()
follower = None
currentleader = connection.getsockname()[0]
skipaddr = connection.getpeername()[0]
myname = get_myname()
for member in cfm.list_collective():
dronecandidate = cfm.get_collective_member(member)['address']
if dronecandidate in (currentleader, skipaddr) or member == myname:
continue
eventlet.spawn_n(try_assimilate, dronecandidate)
def startup():
global cfginitlock
members = list(cfm.list_collective())
if len(members) < 2:
# Not in collective mode, return
return
if cfginitlock is None:
cfginitlock = threading.RLock()
eventlet.spawn_n(start_collective)
def start_collective():
global follower
global retrythread
if follower:
follower.kill()
follower = None
if leader_init.active: # do not start trying to connect if we are
# xmitting data to a follower
return
myname = get_myname()
for member in sorted(list(cfm.list_collective())):
if member == myname:
continue
if cfm.cfgleader is None:
cfm.stop_following(True)
ldrcandidate = cfm.get_collective_member(member)['address']
if connect_to_leader(name=myname, leader=ldrcandidate):
break
else:
retrythread = eventlet.spawn_after(30 + random.random(),
start_collective)
@@ -148,6 +148,14 @@ node = {
# 'autonode.servername, so that would not need to be '
# 'copied ')
# },
'collective.manager': {
'description': ('When in collective mode, the member of the '
'collective currently considered to be responsible '
'for this node. At a future date, this may be '
'modified automatically if another attribute '
'indicates candidate managers, either for '
'high availability or load balancing purposes.')
},
'discovery.policy': {
'description': 'Policy to use for auto-configuration of discovered '
'and identified nodes. Valid values are "manual", '
File diff suppressed because it is too large Load Diff
+170 -9
View File
@@ -23,14 +23,18 @@
# there should be no more than one handler per node
import codecs
import collections
import confluent.collective.manager as collective
import confluent.config.configmanager as configmodule
import confluent.exceptions as exc
import confluent.interface.console as conapi
import confluent.log as log
import confluent.core as plugin
import confluent.tlvdata as tlvdata
import confluent.util as util
import eventlet
import eventlet.event
import eventlet.green.socket as socket
import eventlet.green.ssl as ssl
import pyte
import random
import time
@@ -138,11 +142,13 @@ def pytechars2line(chars, maxlen=None):
class ConsoleHandler(object):
_plugin_path = '/nodes/{0}/_console/session'
_logtobuffer = True
_genwatchattribs = frozenset(('console.method', 'console.logging'))
_genwatchattribs = frozenset(('console.method', 'console.logging',
'collective.manager'))
def __init__(self, node, configmanager):
self.clearpending = False
self._dologging = True
self._is_local = True
self._isondemand = False
self.error = None
self._retrytime = 0
@@ -214,7 +220,7 @@ class ConsoleHandler(object):
def check_isondemand(self):
self._dologging = True
attrvalue = self.cfgmgr.get_node_attributes(
(self.node,), ('console.logging',))
(self.node,), ('console.logging', 'collective.manager'))
if self.node not in attrvalue:
self._isondemand = False
elif 'console.logging' not in attrvalue[self.node]:
@@ -225,6 +231,23 @@ class ConsoleHandler(object):
self._isondemand = True
if (attrvalue[self.node]['console.logging']['value']) in ('none', 'memory'):
self._dologging = False
self.check_collective(attrvalue)
def check_collective(self, attrvalue):
myc = attrvalue.get(self.node, {}).get('collective.manager', {}).get(
'value', None)
if configmodule.list_collective() and not myc:
self._is_local = False
self._detach()
self._disconnect()
if myc and myc != collective.get_myname():
# Do not do console connect for nodes managed by another
# confluent collective member
self._is_local = False
self._detach()
self._disconnect()
else:
self._is_local = True
def get_buffer_age(self):
"""Return age of buffered data
@@ -236,6 +259,10 @@ class ConsoleHandler(object):
return False
def _attribschanged(self, nodeattribs, configmanager, **kwargs):
if 'collective.manager' in nodeattribs[self.node]:
attrval = configmanager.get_node_attributes(self.node,
'collective.manager')
self.check_collective(attrval)
if 'console.logging' in nodeattribs[self.node]:
# decide whether logging changes how we react or not
self._dologging = True
@@ -284,6 +311,10 @@ class ConsoleHandler(object):
'none or interactive,\r\nconnection loss, or service restart]')
self.clearpending = True
def _detach(self):
for ses in list(self.livesessions):
ses.detach()
def _disconnect(self):
if self.connectionthread:
self.connectionthread.kill()
@@ -305,6 +336,8 @@ class ConsoleHandler(object):
self._disconnect()
def _connect(self):
if not self._is_local:
return
if self.connectionthread:
self.connectionthread.kill()
self.connectionthread = None
@@ -320,9 +353,9 @@ class ConsoleHandler(object):
self.reconnect.cancel()
self.reconnect = None
try:
self._console = plugin.handle_path(
self._console = list(plugin.handle_path(
self._plugin_path.format(self.node),
"create", self.cfgmgr)
"create", self.cfgmgr))[0]
except (exc.NotImplementedException, exc.NotFoundException):
self._console = None
except:
@@ -391,7 +424,6 @@ class ConsoleHandler(object):
self._send_rcpts({'connectstate': self.connectstate})
def _got_disconnected(self):
self.clearbuffer()
if self.connectstate != 'unconnected':
self.connectstate = 'unconnected'
self.log(
@@ -400,6 +432,8 @@ class ConsoleHandler(object):
self._send_rcpts({'connectstate': self.connectstate})
if self._isalive:
self._connect()
else:
self.clearbuffer()
def close(self):
self._isalive = False
@@ -412,6 +446,9 @@ class ConsoleHandler(object):
if self.connectionthread:
self.connectionthread.kill()
self.connectionthread = None
if self._attribwatcher:
self.cfgmgr.remove_watcher(self._attribwatcher)
self._attribwatcher = None
def get_console_output(self, data):
# Spawn as a greenthread, return control as soon as possible
@@ -481,9 +518,6 @@ class ConsoleHandler(object):
eventdata |= 1
if self.shiftin is not None:
eventdata |= 2
self.log(data, eventdata=eventdata)
self.lasttime = util.monotonic_time()
self.feedbuffer(data)
# TODO: analyze buffer for registered events, examples:
# panics
# certificate signing request
@@ -492,6 +526,10 @@ class ConsoleHandler(object):
self.feedbuffer(b'\x1bc')
self._send_rcpts(b'\x1bc')
self._send_rcpts(_utf8_normalize(data, self.shiftin, self.utf8decoder))
self.log(data, eventdata=eventdata)
self.lasttime = util.monotonic_time()
self.feedbuffer(data)
def _send_rcpts(self, data):
for rcpt in list(self.livesessions):
@@ -567,7 +605,12 @@ def _nodechange(added, deleting, configmanager):
def _start_tenant_sessions(cfm):
for node in cfm.list_nodes():
nodeattrs = cfm.get_node_attributes(cfm.list_nodes(), 'collective.manager')
for node in nodeattrs:
manager = nodeattrs[node].get('collective.manager', {}).get('value',
None)
if manager and collective.get_myname() != manager:
continue
try:
connect_node(node, cfm)
except:
@@ -583,11 +626,117 @@ def start_console_sessions():
def connect_node(node, configmanager, username=None):
attrval = configmanager.get_node_attributes(node, 'collective.manager')
myc = attrval.get(node, {}).get('collective.manager', {}).get(
'value', None)
myname = collective.get_myname()
if myc and myc != collective.get_myname():
minfo = configmodule.get_collective_member(myc)
return ProxyConsole(node, minfo, myname, configmanager, username)
consk = (node, configmanager.tenant)
if consk not in _handled_consoles:
_handled_consoles[consk] = ConsoleHandler(node, configmanager)
return _handled_consoles[consk]
# A stub console handler that just passes through to a remote confluent
# collective member. It can skip the multi-session sharing as that is handled
# remotely
class ProxyConsole(object):
_genwatchattribs = frozenset(('collective.manager',))
def __init__(self, node, managerinfo, myname, configmanager, user):
self.skipreplay = True
self.managerinfo = managerinfo
self.myname = myname
self.cfm = configmanager
self.node = node
self.user = user
self.remote = None
self.clisession = None
self._attribwatcher = configmanager.watch_attributes(
(self.node,), self._genwatchattribs, self._attribschanged)
def _attribschanged(self, nodeattribs, configmanager, **kwargs):
if self.clisession:
self.clisession.detach()
self.clisession = None
def relay_data(self):
data = tlvdata.recv(self.remote)
while data:
self.data_handler(data)
data = tlvdata.recv(self.remote)
def get_buffer_age(self):
# the server sends a buffer age if appropriate, no need to handle
# it explicitly in the proxy instance
return False
def get_recent(self):
# Again, delegate this to the remote collective member
self.skipreplay = False
return b''
def write(self, data):
# Relay data to the collective manager
try:
tlvdata.send(self.remote, data)
except Exception:
if self.clisession:
self.clisession.detach()
self.clisession = None
def attachsession(self, session):
self.clisession = session
self.data_handler = session.data_handler
termreq = {
'proxyconsole': {
'name': self.myname,
'user': self.user,
'tenant': self.cfm.tenant,
'node': self.node,
'skipreplay': self.skipreplay,
#TODO(jjohnson2): declare myself as a proxy,
#facilitate redirect rather than relay on manager change
},
}
try:
remote = socket.create_connection((self.managerinfo['address'], 13001))
remote = ssl.wrap_socket(remote, cert_reqs=ssl.CERT_NONE,
keyfile='/etc/confluent/privkey.pem',
certfile='/etc/confluent/srvcert.pem')
if not util.cert_matches(self.managerinfo['fingerprint'],
remote.getpeercert(binary_form=True)):
raise Exception('Invalid peer certificate')
except Exception:
eventlet.sleep(3)
if self.clisession:
self.clisession.detach()
self.detachsession(None)
return
tlvdata.recv(remote)
tlvdata.recv(remote)
tlvdata.send(remote, termreq)
self.remote = remote
eventlet.spawn(self.relay_data)
def detachsession(self, session):
# we will disappear, so just let that happen...
if self.remote:
try:
tlvdata.send(self.remote, {'operation': 'stop'})
except Exception:
pass
self.clisession = None
def send_break(self):
tlvdata.send(self.remote, {'operation': 'break'})
def reopen(self):
tlvdata.send(self.remote, {'operation': 'reopen'})
# this represents some api view of a console handler. This handles things like
# holding the caller specific queue data, for example, when http api should be
# sending data, but there is no outstanding POST request to hold it,
@@ -679,6 +828,18 @@ class ConsoleSession(object):
self._evt = None
self.reghdl = None
def detach(self):
"""Handler for the console handler to detach so it can reattach,
currently to facilitate changing from one collective.manager to
another
:return:
"""
self.conshdl.detachsession(self)
self.connect_session()
self.conshdl.attachsession(self)
self.write = self.conshdl.write
def got_data(self, data):
"""Receive data from console and buffer
+220 -15
View File
@@ -35,7 +35,10 @@
import confluent
import confluent.alerts as alerts
import confluent.tlvdata as tlvdata
import confluent.config.attributes as attrscheme
import confluent.config.configmanager as cfm
import confluent.collective.manager as collective
import confluent.discovery.core as disco
import confluent.interface.console as console
import confluent.exceptions as exc
@@ -46,11 +49,27 @@ try:
import confluent.shellmodule as shellmodule
except ImportError:
pass
try:
import OpenSSL.crypto as crypto
except ImportError:
# Only required for collective mode
crypto = None
import confluent.util as util
import eventlet.greenpool as greenpool
import eventlet.green.ssl as ssl
import eventlet.queue as queue
import itertools
import os
try:
import cPickle as pickle
except ImportError:
import pickle
import socket
import struct
import sys
pluginmap = {}
dispatch_plugins = (b'ipmi', u'ipmi')
def seek_element(currplace, currkey):
@@ -154,6 +173,10 @@ def _init_core():
'pluginattrs': ['hardwaremanagement.method'],
'default': 'ipmi',
}),
'hostname': PluginRoute({
'pluginattrs': ['hardwaremanagement.method'],
'default': 'ipmi',
}),
'identifier': PluginRoute({
'pluginattrs': ['hardwaremanagement.method'],
'default': 'ipmi',
@@ -554,6 +577,63 @@ def abbreviate_noderange(configmanager, inputdata, operation):
return (msg.KeyValueData({'noderange': noderange.ReverseNodeRange(inputdata['nodes'], configmanager).noderange}),)
def handle_dispatch(connection, cert, dispatch, peername):
cert = crypto.dump_certificate(crypto.FILETYPE_ASN1, cert)
if not util.cert_matches(
cfm.get_collective_member(peername)['fingerprint'], cert):
connection.close()
return
dispatch = pickle.loads(dispatch)
configmanager = cfm.ConfigManager(dispatch['tenant'])
nodes = dispatch['nodes']
inputdata = dispatch['inputdata']
operation = dispatch['operation']
pathcomponents = dispatch['path']
routespec = nested_lookup(noderesources, pathcomponents)
plugroute = routespec.routeinfo
plugpath = None
nodesbyhandler = {}
passvalues = []
nodeattr = configmanager.get_node_attributes(
nodes, plugroute['pluginattrs'])
for node in nodes:
for attrname in plugroute['pluginattrs']:
if attrname in nodeattr[node]:
plugpath = nodeattr[node][attrname]['value']
elif 'default' in plugroute:
plugpath = plugroute['default']
if plugpath is not None:
try:
hfunc = getattr(pluginmap[plugpath], operation)
except KeyError:
nodesbyhandler[BadPlugin(node, plugpath).error] = [node]
continue
if hfunc in nodesbyhandler:
nodesbyhandler[hfunc].append(node)
else:
nodesbyhandler[hfunc] = [node]
try:
for hfunc in nodesbyhandler:
passvalues.append(hfunc(
nodes=nodesbyhandler[hfunc], element=pathcomponents,
configmanager=configmanager,
inputdata=inputdata))
for res in itertools.chain(*passvalues):
_forward_rsp(connection, res)
except Exception as res:
_forward_rsp(connection, res)
connection.sendall('\x00\x00\x00\x00\x00\x00\x00\x00')
def _forward_rsp(connection, res):
r = pickle.dumps(res)
rlen = len(r)
if not rlen:
return
connection.sendall(struct.pack('!Q', rlen))
connection.sendall(r)
def handle_node_request(configmanager, inputdata, operation,
pathcomponents, autostrip=True):
iscollection = False
@@ -633,7 +713,7 @@ def handle_node_request(configmanager, inputdata, operation,
else:
raise Exception("TODO here")
del pathcomponents[0:2]
passvalues = []
passvalues = queue.Queue()
plugroute = routespec.routeinfo
inputdata = msg.get_input_message(
pathcomponents, operation, inputdata, nodes, isnoderange,
@@ -650,20 +730,35 @@ def handle_node_request(configmanager, inputdata, operation,
if isnoderange:
return passvalue
elif isinstance(passvalue, console.Console):
return passvalue
return [passvalue]
else:
return stripnode(passvalue, nodes[0])
elif 'pluginattrs' in plugroute:
nodeattr = configmanager.get_node_attributes(
nodes, plugroute['pluginattrs'])
nodes, plugroute['pluginattrs'] + ['collective.manager'])
plugpath = None
if 'default' in plugroute:
plugpath = plugroute['default']
nodesbymanager = {}
nodesbyhandler = {}
badcollnodes = []
for node in nodes:
for attrname in plugroute['pluginattrs']:
if attrname in nodeattr[node]:
plugpath = nodeattr[node][attrname]['value']
elif 'default' in plugroute:
plugpath = plugroute['default']
if plugpath in dispatch_plugins:
cfm.check_quorum()
manager = nodeattr[node].get('collective.manager', {}).get(
'value', None)
if manager:
if collective.get_myname() != manager:
if manager not in nodesbymanager:
nodesbymanager[manager] = set([node])
else:
nodesbymanager[manager].add(node)
continue
elif list(cfm.list_collective()):
badcollnodes.append(node)
if plugpath is not None:
try:
hfunc = getattr(pluginmap[plugpath], operation)
@@ -674,19 +769,30 @@ def handle_node_request(configmanager, inputdata, operation,
nodesbyhandler[hfunc].append(node)
else:
nodesbyhandler[hfunc] = [node]
if badcollnodes:
raise exc.ConfluentException(
'collective management active, '
'collective.manager must be set for {0}'.format(
','.join(badcollnodes)))
workers = greenpool.GreenPool()
numworkers = 0
for hfunc in nodesbyhandler:
passvalues.append(hfunc(
nodes=nodesbyhandler[hfunc], element=pathcomponents,
configmanager=configmanager,
inputdata=inputdata))
numworkers += 1
workers.spawn(addtoqueue, passvalues, hfunc, {'nodes': nodesbyhandler[hfunc],
'element': pathcomponents,
'configmanager': configmanager,
'inputdata': inputdata})
for manager in nodesbymanager:
numworkers += 1
workers.spawn(addtoqueue, passvalues, dispatch_request, {
'nodes': nodesbymanager[manager], 'manager': manager,
'element': pathcomponents, 'configmanager': configmanager,
'inputdata': inputdata, 'operation': operation})
if isnoderange or not autostrip:
return itertools.chain(*passvalues)
return iterate_queue(numworkers, passvalues)
else:
if len(passvalues) > 0:
if isinstance(passvalues[0], console.Console):
return passvalues[0]
else:
return stripnode(passvalues[0], nodes[0])
if numworkers > 0:
return iterate_queue(numworkers, passvalues, nodes[0])
else:
raise exc.NotImplementedException()
@@ -696,6 +802,105 @@ def handle_node_request(configmanager, inputdata, operation,
# return stripnode(passvalues[0], nodes[0])
def iterate_queue(numworkers, passvalues, strip=False):
completions = 0
while completions < numworkers:
nv = passvalues.get()
if nv == 'theend':
completions += 1
else:
if strip and not isinstance(nv, console.Console):
nv.strip_node(strip)
yield nv
def addtoqueue(theq, fun, kwargs):
try:
result = fun(**kwargs)
if isinstance(result, console.Console):
theq.put(result)
else:
for pv in result:
theq.put(pv)
finally:
theq.put('theend')
def dispatch_request(nodes, manager, element, configmanager, inputdata,
operation):
a = configmanager.get_collective_member(manager)
try:
remote = socket.create_connection((a['address'], 13001))
remote.settimeout(90)
remote = ssl.wrap_socket(remote, cert_reqs=ssl.CERT_NONE,
keyfile='/etc/confluent/privkey.pem',
certfile='/etc/confluent/srvcert.pem')
except Exception:
for node in nodes:
yield msg.ConfluentResourceUnavailable(
node, 'Collective member {0} is unreachable'.format(a['name']))
return
if not util.cert_matches(a['fingerprint'], remote.getpeercert(
binary_form=True)):
raise Exception("Invalid certificate on peer")
tlvdata.recv(remote)
tlvdata.recv(remote)
myname = collective.get_myname()
dreq = pickle.dumps({'name': myname, 'nodes': list(nodes),
'path': element,'tenant': configmanager.tenant,
'operation': operation, 'inputdata': inputdata})
tlvdata.send(remote, {'dispatch': {'name': myname, 'length': len(dreq)}})
remote.sendall(dreq)
while True:
try:
rlen = remote.recv(8)
except Exception:
for node in nodes:
yield msg.ConfluentResourceUnavailable(
node, 'Collective member {0} went unreachable'.format(
a['name']))
return
while len(rlen) < 8:
try:
nlen = remote.recv(8 - len(rlen))
except Exception:
nlen = 0
if not nlen:
for node in nodes:
yield msg.ConfluentResourceUnavailable(
node, 'Collective member {0} went unreachable'.format(
a['name']))
return
rlen += nlen
rlen = struct.unpack('!Q', rlen)[0]
if rlen == 0:
break
try:
rsp = remote.recv(rlen)
except Exception:
for node in nodes:
yield msg.ConfluentResourceUnavailable(
node, 'Collective member {0} went unreachable'.format(
a['name']))
return
while len(rsp) < rlen:
try:
nrsp = remote.recv(rlen - len(rsp))
except Exception:
nrsp = 0
if not nrsp:
for node in nodes:
yield msg.ConfluentResourceUnavailable(
node, 'Collective member {0} went unreachable'.format(
a['name']))
return
rsp += nrsp
rsp = pickle.loads(rsp)
if isinstance(rsp, Exception):
raise rsp
yield rsp
def handle_discovery(pathcomponents, operation, configmanager, inputdata):
if pathcomponents[0] == 'detected':
pass
+64 -30
View File
@@ -85,6 +85,7 @@ import eventlet
import eventlet.greenpool
import eventlet.semaphore
autosensors = set()
class nesteddict(dict):
def __missing__(self, key):
@@ -350,7 +351,28 @@ def _parameterize_path(pathcomponents):
return validselectors, keyparams, listrequested, childcoll
def handle_autosense_config(operation, inputdata):
autosense = cfm.get_global('discovery.autosense')
autosense = autosense or autosense is None
if operation == 'retrieve':
yield msg.KeyValueData({'enabled': autosense})
elif operation == 'update':
enabled = inputdata['enabled']
if type(enabled) in (unicode, str):
enabled = enabled.lower() in ('true', '1', 'y', 'yes', 'enable',
'enabled')
if autosense == enabled:
return
cfm.set_global('discovery.autosense', enabled)
if enabled:
start_autosense()
else:
stop_autosense()
def handle_api_request(configmanager, inputdata, operation, pathcomponents):
if pathcomponents == ['discovery', 'autosense']:
return handle_autosense_config(operation, inputdata)
if operation == 'retrieve':
return handle_read_api_request(pathcomponents)
elif (operation in ('update', 'create') and
@@ -398,6 +420,7 @@ def handle_read_api_request(pathcomponents):
if len(pathcomponents) == 1:
dirlist = [msg.ChildCollection(x + '/') for x in sorted(list(subcats))]
dirlist.append(msg.ChildCollection('rescan'))
dirlist.append(msg.ChildCollection('autosense'))
return dirlist
if not coll:
return show_info(queryparms['by-mac'])
@@ -695,6 +718,9 @@ def get_chained_smm_name(nodename, cfg, handler, nl=None, checkswitch=True):
smmaddr = cd[nodename]['hardwaremanagement.manager']['value']
pkey = cd[nodename].get('pubkeys.tls_hardwaremanager', {}).get(
'value', None)
if not pkey:
# We cannot continue through a break in the chain
return None, False
if pkey:
cv = util.TLSCertVerifier(
cfg, nodename, 'pubkeys.tls_hardwaremanager').verify_cert
@@ -715,6 +741,8 @@ def get_smm_neighbor_fingerprints(smmaddr, cv):
smmaddr = '[{0}]'.format(smmaddr)
wc = webclient.SecureHTTPConnection(smmaddr, verifycallback=cv)
neighs = wc.grab_json_response('/scripts/neighdata.json')
if not neighs:
return
for idx in (4, 5):
if 'sha256' not in neighs[idx]:
continue
@@ -823,12 +851,7 @@ def eval_node(cfg, handler, info, nodename, manual=False):
handler.probe() # unicast interrogation as possible to get more data
# switch concurrently
# do some preconfig, for example, to bring a SMM online if applicable
errorstr = handler.preconfig()
if errorstr:
if manual:
raise exc.InvalidArgumentException(errorstr)
log.log({'error': errorstr})
return
handler.preconfig()
except Exception as e:
unknown_info[info['hwaddr']] = info
info['discostatus'] = 'unidentified'
@@ -872,8 +895,9 @@ def eval_node(cfg, handler, info, nodename, manual=False):
if enl:
# ambiguous SMM situation according to the configuration, we need
# to match uuid
encuuid = info['attributes'].get('chasis-uuid', None)
encuuid = info['attributes'].get('chassis-uuid', None)
if encuuid:
encuuid = encuuid[0]
enl = list(cfg.filter_node_attributes('id.uuid=' + encuuid))
if len(enl) != 1:
# errorstr = 'No SMM by given UUID known, *yet*'
@@ -929,26 +953,25 @@ def eval_node(cfg, handler, info, nodename, manual=False):
# but... is this really ok? could be on an upstream port or
# erroneously put in the enclosure with no nodes yet
# so first, see if the candidate node is a chain host
if info['maccount']:
# discovery happened through switch
nl = list(cfg.filter_node_attributes(
'enclosure.extends=' + nodename))
if nl:
# The candidate nodename is the head of a chain, we must
# validate the smm certificate by the switch
macmap.get_node_fingerprint(nodename, cfg)
util.handler.cert_matches(fprint, handler.https_cert)
if not manual:
if info.get('maccount', False):
# discovery happened through switch
nl = list(cfg.filter_node_attributes(
'enclosure.extends=' + nodename))
if nl:
# The candidate nodename is the head of a chain, we must
# validate the smm certificate by the switch
macmap.get_node_fingerprint(nodename, cfg)
util.handler.cert_matches(fprint, handler.https_cert)
return
if (info.get('maccount', False) and
not handler.discoverable_by_switch(info['maccount'])):
errorstr = 'The detected node {0} was detected using switch, ' \
'however the relevant port has too many macs learned ' \
'for this type of device ({1}) to be discovered by ' \
'switch.'.format(nodename, handler.devname)
log.log({'error': errorstr})
return
if (info['maccount'] and
not handler.discoverable_by_switch(info['maccount'])):
errorstr = 'The detected node {0} was detected using switch, ' \
'however the relevant port has too many macs learned ' \
'for this type of device ({1}) to be discovered by ' \
'switch.'.format(nodename, handler.devname)
if manual:
raise exc.InvalidArgumentException(errorstr)
log.log({'error': errorstr})
return
if not discover_node(cfg, handler, info, nodename, manual):
pending_nodes[nodename] = info
@@ -1080,12 +1103,14 @@ def newnodes(added, deleting, configmanager):
global attribwatcher
global needaddhandled
global nodeaddhandler
_map_unique_ids()
configmanager.remove_watcher(attribwatcher)
allnodes = configmanager.list_nodes()
attribwatcher = configmanager.watch_attributes(
allnodes, ('discovery.policy', 'net*.switch',
'hardwaremanagement.manager', 'net*.switchport', 'id.uuid',
'pubkeys.tls_hardwaremanager', 'net*.bootable'), _recheck_nodes)
'hardwaremanagement.manager', 'net*.switchport',
'id.uuid', 'pubkeys.tls_hardwaremanager',
'net*.bootable'), _recheck_nodes)
if nodeaddhandler:
needaddhandled = True
else:
@@ -1131,14 +1156,23 @@ def start_detection():
'hardwaremanagement.manager', 'net*.switchport', 'id.uuid',
'pubkeys.tls_hardwaremanager'), _recheck_nodes)
cfg.watch_nodecollection(newnodes)
eventlet.spawn_n(slp.snoop, safe_detected)
eventlet.spawn_n(pxe.snoop, safe_detected)
autosense = cfm.get_global('discovery.autosense')
if autosense or autosense is None:
start_autosense()
if rechecker is None:
rechecktime = util.monotonic_time() + 900
rechecker = eventlet.spawn_after(900, _periodic_recheck, cfg)
# eventlet.spawn_n(ssdp.snoop, safe_detected)
def stop_autosense():
for watcher in list(autosensors):
watcher.kill()
autosensors.discard(watcher)
def start_autosense():
autosensors.add(eventlet.spawn(slp.snoop, safe_detected))
autosensors.add(eventlet.spawn(pxe.snoop, safe_detected))
nodes_by_fprint = {}
@@ -22,24 +22,7 @@ class NodeHandler(immhandler.NodeHandler):
devname = 'XCC'
def preconfig(self):
ff = None
try:
ff = self.info['attributes']['enclosure-form-factor']
except KeyError:
try:
# an XCC should always have that set, this is sign of
# a bug, try to reset the BMC as a workaround
ipmicmd = self._get_ipmicmd()
ipmicmd.reset_bmc()
return "XCC with address {0} did not have attributes " \
"declared, attempting to correct with " \
"XCC reset".format(self.ipaddr)
except pygexc.IpmiException as e:
if (e.ipmicode != 193 and
'Unauthorized name' not in str(e) and
'Incorrect password' not in str(e)):
# raise an issue if anything other than to be expected
raise
ff = self.info.get('attributes', {}).get('enclosure-form-factor', '')
if ff not in ('dense-computing', [u'dense-computing']):
return
# attempt to enable SMM
@@ -16,13 +16,13 @@
import confluent.neighutil as neighutil
import confluent.util as util
import confluent.log as log
import os
import random
import eventlet.green.select as select
import eventlet.green.socket as socket
import struct
import subprocess
import traceback
_slp_services = set([
'service:management-hardware.IBM:integrated-management-module2',
@@ -391,6 +391,7 @@ def snoop(handler):
:param handler:
:return:
"""
tracelog = log.Logger('trace')
active_scan(handler)
net = socket.socket(socket.AF_INET6, socket.SOCK_DGRAM)
net.setsockopt(IPPROTO_IPV6, socket.IPV6_V6ONLY, 1)
@@ -420,54 +421,58 @@ def snoop(handler):
net4.bind(('', 427))
while True:
newmacs = set([])
r, _, _ = select.select((net, net4), (), (), 60)
# clear known_peers and peerbymacaddress
# to avoid stale info getting in...
# rely upon the select(0.2) to catch rapid fire and aggregate ip
# addresses that come close together
# calling code needs to understand deeper context, as snoop
# will now yield dupe info over time
known_peers = set([])
peerbymacaddress = {}
neighutil.update_neigh()
while r:
for s in r:
(rsp, peer) = s.recvfrom(9000)
ip = peer[0].partition('%')[0]
if ip not in neighutil.neightable:
continue
if peer in known_peers:
continue
known_peers.add(peer)
mac = neighutil.neightable[ip]
if mac in peerbymacaddress:
peerbymacaddress[mac]['addresses'].append(peer)
else:
q = query_srvtypes(peer)
if not q or not q[0]:
# SLP might have started and not ready yet
# ignore for now
known_peers.discard(peer)
try:
newmacs = set([])
r, _, _ = select.select((net, net4), (), (), 60)
# clear known_peers and peerbymacaddress
# to avoid stale info getting in...
# rely upon the select(0.2) to catch rapid fire and aggregate ip
# addresses that come close together
# calling code needs to understand deeper context, as snoop
# will now yield dupe info over time
known_peers = set([])
peerbymacaddress = {}
neighutil.update_neigh()
while r:
for s in r:
(rsp, peer) = s.recvfrom(9000)
ip = peer[0].partition('%')[0]
if ip not in neighutil.neightable:
continue
# we want to prioritize the very well known services
svcs = []
for svc in q:
if svc in _slp_services:
svcs.insert(0, svc)
else:
svcs.append(svc)
peerbymacaddress[mac] = {
'services': svcs,
'addresses': [peer],
}
newmacs.add(mac)
r, _, _ = select.select((net, net4), (), (), 0.2)
for mac in newmacs:
peerbymacaddress[mac]['xid'] = 1
_add_attributes(peerbymacaddress[mac])
peerbymacaddress[mac]['hwaddr'] = mac
handler(peerbymacaddress[mac])
if peer in known_peers:
continue
known_peers.add(peer)
mac = neighutil.neightable[ip]
if mac in peerbymacaddress:
peerbymacaddress[mac]['addresses'].append(peer)
else:
q = query_srvtypes(peer)
if not q or not q[0]:
# SLP might have started and not ready yet
# ignore for now
known_peers.discard(peer)
continue
# we want to prioritize the very well known services
svcs = []
for svc in q:
if svc in _slp_services:
svcs.insert(0, svc)
else:
svcs.append(svc)
peerbymacaddress[mac] = {
'services': svcs,
'addresses': [peer],
}
newmacs.add(mac)
r, _, _ = select.select((net, net4), (), (), 0.2)
for mac in newmacs:
peerbymacaddress[mac]['xid'] = 1
_add_attributes(peerbymacaddress[mac])
peerbymacaddress[mac]['hwaddr'] = mac
handler(peerbymacaddress[mac])
except Exception as e:
tracelog.log(traceback.format_exc(), ltype=log.DataTypes.event,
event=log.Events.stacktrace)
def active_scan(handler):
@@ -191,7 +191,6 @@ def _parse_ssdp(peer, rsp, peerdata):
_, code, _ = headlines[0].split(' ', 2)
except ValueError:
return
myurl = None
if code == '200':
if nid in peerdata:
peerdatum = peerdata[nid]
@@ -208,7 +207,6 @@ def _parse_ssdp(peer, rsp, peerdata):
header = header.strip()
value = value.strip()
if header == 'AL' or header == 'LOCATION':
myurl = value
if 'urls' not in peerdatum:
peerdatum['urls'] = [value]
elif value not in peerdatum['urls']:
+5
View File
@@ -68,6 +68,11 @@ class LockedCredentials(ConfluentException):
_apierrorstr = 'Credential store locked'
class DegradedCollective(ConfluentException):
# We are in a collective with at least half of the member nodes missing
_apierrorstr = 'Collective does not have quorum'
class ForbiddenRequest(ConfluentException):
# The client request is not allowed by authorization engine
apierrorcode = 403
@@ -21,6 +21,8 @@
import confluent.exceptions as exc
import confluent.messages as msg
import eventlet
import os
import socket
updatesbytarget = {}
uploadsbytarget = {}
@@ -28,6 +30,11 @@ updatepool = eventlet.greenpool.GreenPool(256)
def execupdate(handler, filename, updateobj, type):
if not os.path.exists(filename):
errstr = '{0} does not appear to exist on {1}'.format(
filename, socket.gethostname())
updateobj.handle_progress({'phase': 'error', 'progress': 0.0,
'detail': errstr})
try:
if type == 'firmware':
completion = handler(filename, progress=updateobj.handle_progress,
+6 -3
View File
@@ -400,7 +400,7 @@ def resourcehandler_backend(env, start_response):
('Pragma', 'no-cache'),
('X-Content-Type-Options', 'nosniff'),
('Content-Security-Policy', "default-src 'self'"),
('X-XSS-Protection', '1'), ('X-Frame-Options', 'deny'),
('X-XSS-Protection', '1; mode=block'), ('X-Frame-Options', 'deny'),
('Strict-Transport-Security', 'max-age=86400'),
('X-Permitted-Cross-Domain-Policies', 'none')]
reqbody = None
@@ -741,7 +741,10 @@ def _assemble_json(responses, resource=None, url=None, extension=None):
for dk in rsp.iterkeys():
if dk in rspdata:
if isinstance(rspdata[dk], list):
rspdata[dk].append(rsp[dk])
if isinstance(rsp[dk], list):
rspdata[dk].extend(rsp[dk])
else:
rspdata[dk].append(rsp[dk])
else:
rspdata[dk] = [rspdata[dk], rsp[dk]]
else:
@@ -782,7 +785,7 @@ def serve(bind_host, bind_port):
' a second\n')
eventlet.sleep(1)
eventlet.wsgi.server(sock, resourcehandler, log=False, log_output=False,
debug=False)
debug=False, socket_timeout=60)
class HttpApi(object):
+2
View File
@@ -33,6 +33,7 @@ import confluent.consoleserver as consoleserver
import confluent.core as confluentcore
import confluent.httpapi as httpapi
import confluent.log as log
import confluent.collective.manager as collective
try:
import confluent.sockapi as sockapi
except ImportError:
@@ -228,6 +229,7 @@ def run():
_updatepidfile()
signal.signal(signal.SIGINT, terminate)
signal.signal(signal.SIGTERM, terminate)
collective.startup()
if dbgif:
oumask = os.umask(0077)
try:
+33
View File
@@ -397,6 +397,9 @@ def get_input_message(path, operation, inputdata, nodes=None, multinode=False,
elif (path[:3] == ['configuration', 'management_controller', 'identifier']
and operation != 'retrieve'):
return InputMCI(path, nodes, inputdata)
elif (path[:3] == ['configuration', 'management_controller', 'hostname']
and operation != 'retrieve'):
return InputHostname(path, nodes, inputdata, configmanager)
elif (path[:4] == ['configuration', 'management_controller',
'net_interfaces', 'management'] and operation != 'retrieve'):
return InputNetworkConfiguration(path, nodes, inputdata,
@@ -719,6 +722,25 @@ class InputBMCReset(ConfluentInputMessage):
return self.inputbynode[node]
class InputHostname(ConfluentInputMessage):
def __init__(self, path, nodes, inputdata, configmanager):
self.inputbynode = {}
self.stripped = False
if not inputdata or 'hostname' not in inputdata:
raise exc.InvalidArgumentException('missing hostname attribute')
if nodes is None:
raise exc.InvalidArgumentException(
'This only supports per-node input')
for expanded in configmanager.expand_attrib_expression(
nodes, inputdata['hostname']):
node, value = expanded
self.inputbynode[node] = value
def hostname(self, node):
return self.inputbynode[node]
class InputMCI(ConfluentInputMessage):
def __init__(self, path, nodes, inputdata):
self.inputbynode = {}
@@ -1298,6 +1320,17 @@ class MCI(ConfluentMessage):
self.kvpairs = {name: kv}
class Hostname(ConfluentMessage):
def __init__(self, name=None, hostname=None):
self.notnode = name is None
self.desc = 'BMC hostname'
kv = {'hostname': {'value': hostname}}
if self.notnode:
self.kvpairs = kv
else:
self.kvpairs = {name: kv}
class DomainName(ConfluentMessage):
def __init__(self, name=None, dn=None):
self.notnode = name is None
+26 -1
View File
@@ -16,6 +16,7 @@
# this will implement noderange grammar
import confluent.exceptions as exc
import codecs
import struct
import eventlet.green.socket as socket
@@ -121,4 +122,28 @@ def get_prefix_len_for_ip(ip):
nbits += 1
maskn = maskn << 1 & 0xffffffff
return nbits
raise exc.NotImplementedException("Non local addresses not supported")
raise exc.NotImplementedException("Non local addresses not supported")
def addresses_match(addr1, addr2):
"""Check two network addresses for similarity
Is it zero padded in one place, not zero padded in another? Is one place by name and another by IP??
Is one context getting a normal IPv4 address and another getting IPv4 in IPv6 notation?
This function examines the two given names, performing the required changes to compare them for equivalency
:param addr1:
:param addr2:
:return: True if the given addresses refer to the same thing
"""
for addrinfo in socket.getaddrinfo(addr1, 0, 0, socket.SOCK_STREAM):
rootaddr1 = socket.inet_pton(addrinfo[0], addrinfo[4][0])
if addrinfo[0] == socket.AF_INET6 and rootaddr1[:12] == b'\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\xff\xff':
# normalize to standard IPv4
rootaddr1 = rootaddr1[-4:]
for otherinfo in socket.getaddrinfo(addr2, 0, 0, socket.SOCK_STREAM):
otheraddr = socket.inet_pton(otherinfo[0], otherinfo[4][0])
if otherinfo[0] == socket.AF_INET6 and otheraddr[:12] == b'\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\xff\xff':
otheraddr = otheraddr[-4:]
if otheraddr == rootaddr1:
return True
return False
@@ -88,7 +88,7 @@ _chassisidbyswitch = {}
def lenovoname(idx, desc):
if desc.isdigit():
return 'Ethernet' + str(idx)
return 'Ethernet' + str(desc)
return desc
nameoverrides = [
@@ -151,6 +151,9 @@ def _extract_extended_desc(info, source, integritychecked):
info['peerdescription'] = source
def sanitize(val):
# This is pretty much the same approach net-snmp takes.
# if the string is printable as-is, then just give it as-is
# if the string has non-printable, then hexify it
val = str(val)
for x in val.strip('\x00'):
if ord(x) < 32 or ord(x) > 128:
@@ -161,7 +164,7 @@ def sanitize(val):
def _init_lldp(data, iname, idx, idxtoportid, switch):
if iname not in data:
data[iname] = {'port': iname, 'portid': str(idxtoportid[idx]),
'chassisid': str(_chassisidbyswitch[switch])}
'chassisid': _chassisidbyswitch[switch]}
def _extract_neighbor_data_b(args):
"""Build LLDP data about elements connected to switch
@@ -180,7 +183,8 @@ def _extract_neighbor_data_b(args):
sid = str(sysid[1][6:])
idxtoifname = {}
idxtoportid = {}
_chassisidbyswitch[switch] = list(conn.walk('1.0.8802.1.1.2.1.3.2'))[0][1]
_chassisidbyswitch[switch] = sanitize(list(
conn.walk('1.0.8802.1.1.2.1.3.2'))[0][1])
for oidindex in conn.walk('1.0.8802.1.1.2.1.3.7.1.3'):
idx = oidindex[0][-1]
idxtoportid[idx] = sanitize(oidindex[1])
@@ -462,7 +462,7 @@ def handle_read_api_request(pathcomponents, configmanager):
def dump_macinfo(macaddr):
macaddr = macaddr.replace('-', ':')
macaddr = macaddr.replace('-', ':').lower()
info = _macmap.get(macaddr, None)
if info is None:
raise exc.NotFoundException(
@@ -20,6 +20,7 @@ import confluent.util as util
def retrieve(nodes, element, configmanager, inputdata):
configmanager.check_quorum()
if nodes is not None:
return retrieve_nodes(nodes, element, configmanager, inputdata)
elif element[0] == 'nodegroups':
@@ -183,8 +184,10 @@ def _expand_expression(nodes, configmanager, inputdata):
pernodeexpressions[expanded[0]] = expanded[1]
for node in util.natural_sort(pernodeexpressions):
yield msg.KeyValueData({'value': pernodeexpressions[node]}, node)
except ValueError as e:
raise exc.InvalidArgumentException(str(e))
except (SyntaxError, ValueError) as e:
raise exc.InvalidArgumentException(
'Bad confluent expression syntax (must use "{{" and "}}" if not '
'desiring confluent expansion): ' + str(e))
@@ -132,7 +132,7 @@ class IpmiCommandWrapper(ipmicommand.Command):
self._inhealth = False
self._lasthealth = None
self._attribwatcher = cfm.watch_attributes(
(node,), ('secret.hardwaremanagementuser',
(node,), ('secret.hardwaremanagementuser', 'collective.manager',
'secret.hardwaremanagementpassword', 'secret.ipmikg',
'hardwaremanagement.manager'), self._attribschanged)
super(self.__class__, self).__init__(**kwargs)
@@ -270,13 +270,17 @@ class IpmiConsole(conapi.Console):
kg=self.kg, force=True,
iohandler=self.handle_data)
self.solconnection.outputlock = NullLock()
while not self.solconnection.connected and not self.broken:
while (self.solconnection and not self.solconnection.connected and
not (self.broken or self.solconnection.broken or
self.solconnection.ipmi_session.broken)):
w = eventlet.event.Event()
_ipmiwaiters.append(w)
w.wait()
if self.broken:
break
if self.broken:
w.wait(15)
if (self.broken or not self.solconnection or
self.solconnection.broken or
self.solconnection.ipmi_session.broken):
if not self.error:
self.error = 'Unknown error'
if (self.error.startswith('Incorrect password') or
self.error.startswith('Unauthorized name')):
raise exc.TargetEndpointBadCredentials
@@ -292,6 +296,7 @@ class IpmiConsole(conapi.Console):
self.solconnection.out_handler = _donothing
self.solconnection.close()
self.solconnection = None
self.datacallback = None
self.broken = True
self.error = "closed"
@@ -344,7 +349,7 @@ def perform_request(operator, node, element,
cfg, results).handle_request()
except pygexc.IpmiException as ipmiexc:
excmsg = str(ipmiexc)
if excmsg == 'Session no longer connected':
if excmsg in ('Session no longer connected', 'timeout'):
results.put(msg.ConfluentTargetTimeout(node))
else:
results.put(msg.ConfluentNodeError(node, excmsg))
@@ -390,7 +395,8 @@ class IpmiHandler(object):
self.tenant = cfg.tenant
tenant = cfg.tenant
if ((node, tenant) not in persistent_ipmicmds or
not persistent_ipmicmds[(node, tenant)].ipmi_session.logged):
not persistent_ipmicmds[(node, tenant)].ipmi_session.logged or
persistent_ipmicmds[(node, tenant)].ipmi_session.broken):
try:
persistent_ipmicmds[(node, tenant)].close_confluent()
except KeyError: # was no previous session
@@ -405,9 +411,10 @@ class IpmiHandler(object):
ipmisess = persistent_ipmicmds[(node, tenant)].ipmi_session
begin = util.monotonic_time()
while ((not (self.broken or self.loggedin)) and
(util.monotonic_time() - begin) < 180):
ipmisess.wait_for_rsp(180)
(util.monotonic_time() - begin) < 30):
ipmisess.wait_for_rsp(31 - (util.monotonic_time() - begin))
if not (self.broken or self.loggedin):
ipmisess._mark_broken()
raise exc.TargetEndpointUnreachable(
"Login process to " + connparams['bmc'] + " died")
except socket.gaierror as ge:
@@ -518,6 +525,8 @@ class IpmiHandler(object):
return self.handle_reset()
elif self.element[1:3] == ['management_controller', 'identifier']:
return self.handle_identifier()
elif self.element[1:3] == ['management_controller', 'hostname']:
return self.handle_hostname()
elif self.element[1:3] == ['management_controller', 'domain_name']:
return self.handle_domain_name()
elif self.element[1:3] == ['management_controller', 'ntp']:
@@ -595,10 +604,20 @@ class IpmiHandler(object):
))
elif self.op == 'update':
config = self.inputdata.netconfig(self.node)
self.ipmicmd.set_net_configuration(
ipv4_address=config['ipv4_address'],
ipv4_configuration=config['ipv4_configuration'],
ipv4_gateway=config['ipv4_gateway'])
try:
self.ipmicmd.set_net_configuration(
ipv4_address=config['ipv4_address'],
ipv4_configuration=config['ipv4_configuration'],
ipv4_gateway=config['ipv4_gateway'])
except socket.error as se:
self.output.put(msg.ConfluentNodeError(self.node,
se.message))
except ValueError as e:
if e.message == 'negative shift count':
self.output.put(msg.ConfluentNodeError(
self.node, 'Invalid prefix length given'))
else:
raise
def handle_users(self):
# Create user
@@ -986,6 +1005,16 @@ class IpmiHandler(object):
self.ipmicmd.set_mci(mci)
return
def handle_hostname(self):
if 'read' == self.op:
hostname = self.ipmicmd.get_hostname()
self.output.put(msg.Hostname(self.node, hostname))
return
elif 'update' == self.op:
hostname = self.inputdata.hostname(self.node)
self.ipmicmd.set_hostname(hostname)
return
def handle_domain_name(self):
if 'read' == self.op:
dn = self.ipmicmd.get_domain_name()
@@ -63,6 +63,7 @@ class SshShell(conapi.Console):
def __init__(self, node, config, username='', password=''):
self.node = node
self.ssh = None
self.datacallback = None
self.nodeconfig = config
self.username = username
self.password = password
@@ -128,7 +129,7 @@ class SshShell(conapi.Console):
data = data[:delidx - 1] + data[delidx + 1:]
self.username += data
if '\r' in self.username:
self.username, self.password = self.username.split('\r')
self.username, self.password = self.username.split('\r')[:2]
lastdata = data.split('\r')[0]
if lastdata != '':
self.datacallback(lastdata)
@@ -155,6 +156,7 @@ class SshShell(conapi.Console):
def close(self):
if self.ssh is not None:
self.ssh.close()
self.datacallback = None
def create(nodes, element, configmanager, inputdata):
if len(nodes) == 1:
@@ -116,6 +116,7 @@ class ExecConsole(conapi.Console):
break
if self.subproc is not None and self.subproc.poll() is None:
self.subproc.kill()
self._datacallback = None
class Plugin(object):
@@ -30,6 +30,9 @@ class _ShellHandler(consoleserver.ConsoleHandler):
_genwatchattribs = False
_logtobuffer = False
def check_collective(self, attrvalue):
return
def log(self, *args, **kwargs):
# suppress logging through proving a stub 'log' function
return
+4
View File
@@ -99,6 +99,10 @@ class Session(object):
elif errnum:
raise exc.ConfluentException(errnum.prettyPrint())
for ans in answers:
if not obj[0].isPrefixOf(ans[0]):
# PySNMP returns leftovers in a bulk command
# filter out such leftovers
break
yield ans
except snmperr.WrongValueError:
raise exc.TargetEndpointBadCredentials('Invalid SNMPv3 password')
+142 -13
View File
@@ -21,6 +21,8 @@
#
import atexit
import ctypes
import ctypes.util
import errno
import os
import pwd
@@ -29,6 +31,7 @@ import struct
import sys
import traceback
import eventlet.green.select as select
import eventlet.green.socket as socket
import eventlet.green.ssl as ssl
import eventlet
@@ -41,7 +44,8 @@ import confluent.exceptions as exc
import confluent.log as log
import confluent.core as pluginapi
import confluent.shellserver as shellserver
import confluent.collective.manager as collective
import confluent.util as util
tracelog = None
auditlog = None
@@ -54,6 +58,22 @@ except AttributeError:
else:
SO_PEERCRED = 17
try:
# Python core TLS despite improvements still has no support for custom
# verify functions.... try to use PyOpenSSL where available to support
# client certificates with custom verify
import eventlet.green.OpenSSL.SSL as libssl
# further, not even pyopenssl exposes SSL_CTX_set_cert_verify_callback
# so we need to ffi that in using a strategy compatible with PyOpenSSL
import OpenSSL.SSL as libssln
import OpenSSL.crypto as crypto
from OpenSSL._util import ffi
except ImportError:
libssl = None
ffi = None
crypto = None
plainsocket = None
class ClientConsole(object):
def __init__(self, client):
@@ -82,7 +102,7 @@ def send_data(connection, data):
raise
def sessionhdl(connection, authname, skipauth=False):
def sessionhdl(connection, authname, skipauth=False, cert=None):
# For now, trying to test the console stuff, so let's just do n4.
authenticated = False
authdata = None
@@ -99,6 +119,15 @@ def sessionhdl(connection, authname, skipauth=False):
while not authenticated: # prompt for name and passphrase
send_data(connection, {'authpassed': 0})
response = tlvdata.recv(connection)
if 'collective' in response:
return collective.handle_connection(connection, cert,
response['collective'])
if 'dispatch' in response:
dreq = tlvdata.recvall(connection, response['dispatch']['length'])
return pluginapi.handle_dispatch(connection, cert, dreq,
response['dispatch']['name'])
if 'proxyconsole' in response:
return start_proxy_term(connection, cert, response['proxyconsole'])
authname = response['username']
passphrase = response['password']
# note(jbjohnso): here, we need to authenticate, but not
@@ -114,6 +143,18 @@ def sessionhdl(connection, authname, skipauth=False):
cfm = authdata[1]
send_data(connection, {'authpassed': 1})
request = tlvdata.recv(connection)
if 'collective' in request and skipauth:
if not libssl:
tlvdata.send(
connection,
{'collective': {'error': 'Server either does not have '
'python-pyopenssl installed or has an '
'incorrect version installed '
'(e.g. pyOpenSSL would need to be '
'replaced with python-pyopenssl)'}})
return
return collective.handle_connection(connection, None, request['collective'],
local=True)
while request is not None:
try:
process_request(
@@ -172,7 +213,7 @@ def process_request(connection, request, cfm, authdata, authname, skipauth):
if operation == 'start':
return start_term(authname, cfm, connection, params, path,
authdata, skipauth)
elif operation == 'shutdown':
elif operation == 'shutdown' and skipauth:
configmanager.ConfigManager.shutdown()
else:
hdlr = pluginapi.handle_path(path, operation, cfm, params)
@@ -187,6 +228,18 @@ def process_request(connection, request, cfm, authdata, authname, skipauth):
send_response(hdlr, connection)
return
def start_proxy_term(connection, cert, request):
cert = crypto.dump_certificate(crypto.FILETYPE_ASN1, cert)
droneinfo = configmanager.get_collective_member(request['name'])
if not util.cert_matches(droneinfo['fingerprint'], cert):
connection.close()
return
cfm = configmanager.ConfigManager(request['tenant'])
ccons = ClientConsole(connection)
consession = consoleserver.ConsoleSession(
node=request['node'], configmanager=cfm, username=request['user'],
datacallback=ccons.sendall, skipreplay=request['skipreplay'])
term_interact(None, None, ccons, None, connection, consession, None)
def start_term(authname, cfm, connection, params, path, authdata, skipauth):
elems = path.split('/')
@@ -210,12 +263,16 @@ def start_term(authname, cfm, connection, params, path, authdata, skipauth):
node=node, configmanager=cfm, username=authname,
datacallback=ccons.sendall, skipreplay=skipreplay,
sessionid=sessionid)
else:
raise exc.InvalidArgumentException('Invalid path {0}'.format(path))
if consession is None:
raise Exception("TODO")
term_interact(authdata, authname, ccons, cfm, connection, consession,
skipauth)
def term_interact(authdata, authname, ccons, cfm, connection, consession,
skipauth):
send_data(connection, {'started': 1})
ccons.startsending()
bufferage = consession.get_buffer_age()
@@ -226,7 +283,7 @@ def start_term(authname, cfm, connection, params, path, authdata, skipauth):
if type(data) == dict:
if data['operation'] == 'stop':
consession.destroy()
return
break
elif data['operation'] == 'break':
consession.send_break()
continue
@@ -247,11 +304,12 @@ def start_term(authname, cfm, connection, params, path, authdata, skipauth):
continue
if not data:
consession.destroy()
return
break
consession.write(data)
def _tlshandler(bind_host, bind_port):
global plainsocket
plainsocket = socket.socket(socket.AF_INET6)
plainsocket.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1)
plainsocket.setsockopt(socket.IPPROTO_TCP, socket.TCP_NODELAY, 1)
@@ -271,13 +329,54 @@ def _tlshandler(bind_host, bind_port):
eventlet.spawn_n(_tlsstartup, cnn)
if ffi:
@ffi.callback("int(*)( X509_STORE_CTX *, void*)")
def verify_stub(store, misc):
return 1
def _tlsstartup(cnn):
authname = None
cnn = ssl.wrap_socket(cnn, keyfile="/etc/confluent/privkey.pem",
certfile="/etc/confluent/srvcert.pem",
ssl_version=ssl.PROTOCOL_TLSv1,
server_side=True)
sessionhdl(cnn, authname)
cert = None
if libssl:
# most fully featured SSL function
ctx = libssl.Context(libssl.SSLv23_METHOD)
ctx.set_options(libssl.OP_NO_SSLv2 | libssl.OP_NO_SSLv3 |
libssl.OP_NO_TLSv1 | libssl.OP_NO_TLSv1_1 |
libssl.OP_CIPHER_SERVER_PREFERENCE)
ctx.set_cipher_list(
'ECDHE-RSA-AES128-GCM-SHA256:ECDHE-RSA-AES256-GCM-SHA384:'
'ECDHE-ECDSA-AES128-GCM-SHA256:ECDHE-ECDSA-AES256-GCM-SHA384')
ctx.set_tmp_ecdh(crypto.get_elliptic_curve('secp384r1'))
ctx.use_certificate_file('/etc/confluent/srvcert.pem')
ctx.use_privatekey_file('/etc/confluent/privkey.pem')
ctx.set_verify(libssln.VERIFY_PEER, lambda *args: True)
libssln._lib.SSL_CTX_set_cert_verify_callback(ctx._context,
verify_stub, ffi.NULL)
cnn = libssl.Connection(ctx, cnn)
cnn.set_accept_state()
cnn.do_handshake()
cert = cnn.get_peer_certificate()
else:
try:
# Try relatively newer python TLS function
ctx = ssl.SSLContext(ssl.PROTOCOL_SSLv23)
ctx.options |= ssl.OP_NO_SSLv2 | ssl.OP_NO_SSLv3
ctx.options |= ssl.OP_NO_TLSv1 | ssl.OP_NO_TLSv1_1
ctx.options |= ssl.OP_CIPHER_SERVER_PREFERENCE
ctx.set_ciphers(
'ECDHE-RSA-AES128-GCM-SHA256:ECDHE-RSA-AES256-GCM-SHA384:'
'ECDHE-ECDSA-AES128-GCM-SHA256:ECDHE-ECDSA-AES256-GCM-SHA384')
ctx.load_cert_chain('/etc/confluent/srvcert.pem',
'/etc/confluent/privkey.pem')
cnn = ctx.wrap_socket(cnn, server_side=True)
except AttributeError:
# Python 2.6 era, go with best effort
cnn = ssl.wrap_socket(cnn, keyfile="/etc/confluent/privkey.pem",
certfile="/etc/confluent/srvcert.pem",
ssl_version=ssl.PROTOCOL_TLSv1,
server_side=True)
sessionhdl(cnn, authname, cert=cert)
def removesocket():
try:
@@ -336,6 +435,36 @@ class SockApi(object):
global tracelog
tracelog = log.Logger('trace')
auditlog = log.Logger('audit')
self.tlsserver = None
if self.should_run_remoteapi():
self.start_remoteapi()
else:
eventlet.spawn_n(self.watch_for_cert)
self.unixdomainserver = eventlet.spawn(_unixdomainhandler)
def watch_for_cert(self):
libc = ctypes.CDLL(ctypes.util.find_library('c'))
watcher = libc.inotify_init()
if libc.inotify_add_watch(watcher, '/etc/confluent/', 0x100) > -1:
while True:
select.select((watcher,), (), (), 86400)
if self.should_run_remoteapi():
os.close(watcher)
self.start_remoteapi()
break
def should_run_remoteapi(self):
return os.path.exists("/etc/confluent/srvcert.pem")
def stop_remoteapi(self):
if self.tlsserver is None:
return
self.tlsserver.kill()
plainsocket.close()
self.tlsserver = None
def start_remoteapi(self):
if self.tlsserver is not None:
return
self.tlsserver = eventlet.spawn(
_tlshandler, self.bind_host, self.bind_port)
self.unixdomainserver = eventlet.spawn(_unixdomainhandler)
+15
View File
@@ -24,6 +24,7 @@ import netifaces
import os
import re
import socket
import ssl
import struct
@@ -99,6 +100,20 @@ def monotonic_time():
return os.times()[4]
def get_certificate_from_file(certfile):
cert = open(certfile, 'rb').read()
inpemcert = False
prunedcert = ''
for line in cert.split('\n'):
if '-----BEGIN CERTIFICATE-----' in line:
inpemcert = True
if inpemcert:
prunedcert += line
if '-----END CERTIFICATE-----' in line:
break
return ssl.PEM_cert_to_DER_cert(prunedcert)
def get_fingerprint(certificate, algo='sha512'):
if algo == 'sha256':
return 'sha256$' + hashlib.sha256(certificate).hexdigest()
+1 -1
View File
@@ -12,7 +12,7 @@ Group: Development/Libraries
BuildRoot: %{_tmppath}/%{name}-%{version}-%{release}-buildroot
Prefix: %{_prefix}
BuildArch: noarch
Requires: python-pyghmi >= 1.0.34, python-eventlet, python-greenlet, python-crypto >= 2.6.1, confluent_client, pyparsing, python-paramiko, python-dns, python-netifaces, python2-pyasn1 >= 0.2.3, python-pysnmp >= 4.3.4, python-pyte, python-lxml, python-eficompressor
Requires: python-pyghmi >= 1.0.34, python-eventlet, python-greenlet, python-pycryptodomex >= 3.4.7, confluent_client, python-pyparsing, python-paramiko, python-dns, python-netifaces, python2-pyasn1 >= 0.2.3, python-pysnmp >= 4.3.4, python-pyte, python-lxml, python-eficompressor, python-setuptools
Vendor: Jarrod Johnson <jjohnson2@lenovo.com>
Url: http://xcat.sf.net/
+3 -1
View File
@@ -6,4 +6,6 @@ if [ "$NUMCOMMITS" != "$VERSION" ]; then
fi
echo $VERSION > VERSION
sed -e "s/#VERSION#/$VERSION/" setup.py.tmpl > setup.py
echo '__version__ = "'$VERSION'"' > confluent/__init__.py
if [ -f confluent/client.py ]; then
echo '__version__ = "'$VERSION'"' > confluent/__init__.py
fi
+4 -2
View File
@@ -15,10 +15,12 @@ setup(
'confluent/networking/',
'confluent/plugins/hardwaremanagement/',
'confluent/plugins/shell/',
'confluent/collective/',
'confluent/plugins/configuration/'],
install_requires=['paramiko', 'pycrypto>=2.6', 'confluent_client>=0.1.0', 'eventlet',
'pyghmi>=0.6.5'],
scripts=['bin/confluent', 'bin/confluentdbutil'],
'dnspython', 'netifaces', 'pyte', 'pysnmp', 'pyparsing',
'pyghmi>=1.0.44'],
scripts=['bin/confluent', 'bin/confluentdbutil', 'bin/collective'],
data_files=[('/etc/init.d', ['sysvinit/confluent']),
('/usr/lib/sysctl.d', ['sysctl/confluent.conf']),
('/usr/lib/systemd/system', ['systemd/confluent.service']),